Article 77SYW Secret Tracking Device Placed in Rare Book Ends Up in Amazon Processing Facility

Secret Tracking Device Placed in Rare Book Ends Up in Amazon Processing Facility

by
jelizondo
from SoylentNews on (#77SYW)

Arthur T Knackerbracket writes:

Destroying books to train AI models is 'all' the Vegas warehouse does:

An operation exposed by an AirTag

Employees who spoke to 404 Media said they spend their time at the facility cutting the spines off of books and feeding them into a scanner. VGT3 is reportedly part of another, larger facility in Las Vegas called LAS8. Its purpose, according to the employees who spoke to 404 media, is solely to cut the spines off books and scan them. "All we do is scan books," one employee told the outlet.

Often, these large orders contain a scattershot of different books. An independent bookseller in Ireland, for example, received an order for 5,000 books that they suspected was for training AI. Some was high-quality non-fiction, like A History of Connemara, and the next thing might be The Eddie Hobbs Guide to your SSIA," Tomas Kenny of Kennys Bookshop said at the time. 404 Media says that employees at VGT3 are required to scan the ISBN (barcode) of each book, lending credibility to the theory that AI companies are working through a list of every book that has ever been published.

Or, at least, every book with an ISBN. A bookseller told 404 Media that these large orders "never" include rare books that don't have an ISBN.

Earlier this year, court filings revealed Anthropic's 'Project Panama,' which kicked off in 2024. Documents as part of those filings make the project's purpose clear: "Project Panama is our effort to destructively scan all the books in the world," said an internal planning document. A judge ruled that Anthropic was allowed to use books to train its AI models. However, it was fined $1.5 billion for keeping 7 million pirated books in a central library. In a 2025 lawsuit against Meta, it was revealed that it had pirated nearly 82TB of books.

Scanning books en masse is nothing new. In 2005, Google spearheaded Google Books by scanning out-of-copyright titles from an on-campus library using specialized scanners. The books were then returned. Presumably, Google made use of some type of V-shaped scanner, which lays a book out naturally so as not to disturb the spine and binding.

In the case of VGT3, 404 reports that employees cut off the spine of the book before scanning, presumably to feed the pages flat into an industrial scanner. The insatiable hunger for data in frontier AI models seems to be moving at a faster pace than Google's early book digitization efforts.

It's hard to say why a facility like VGT3 operates in this way, though it likely comes down to cost. As major companies like Meta and Anthropic have been caught with a library of pirated books, they now need to buy them. And when purchasing thousands of books at a time, it's probably much cheaper to get secondhand copies from marketplaces like Biblio than it is to spend full price on digital versions of those books (if digital versions exist in the first place).

Original Submission

Read more of this story at SoylentNews.

External Content
Source RSS or Atom Feed
Feed Location https://soylentnews.org/index.rss
Feed Title SoylentNews
Feed Link https://soylentnews.org/
Feed Copyright Copyright 2014, SoylentNews
Reply 0 comments