Amazon is buying massive quantities of books, cutting their spines off for faster scanning, and destroying them in the process. A 404 Media investigation tracked a shipment of rare books with an AirTag to Amazon's VGT3 warehouse in Las Vegas, where workers do nothing but receive, cut, scan, and discard. The legal precedent that makes this possible, from a judge ruling on Anthropic's Project Panama, held that destroying the original copy strengthened the fair use defense. This creates a legal architecture where the optimal strategy is to buy and destroy. Meanwhile, books printed before 2022 command a premium because they are guaranteed free of AI-generated text. The contamination of the internet by AI output has made the uncontaminated analog original more valuable than ever. The remedy for digital pollution is to consume the physical world.
404 Media placed an Apple AirTag in a shipment of rare books. The shipment flew from California to Milwaukee, sat in a distribution warehouse near Kenosha for two weeks, then traveled by truck across Colorado to Las Vegas. It arrived at an Amazon warehouse called LAS8. Inside the north end of LAS8 is a separate operation with a different code: VGT3. Its logo, painted on the entrance, is a Tyrannosaurus rex holding an open book.
VGT3 is where Amazon takes delivery of printed books, cuts the spines off, scans the pages, and discards what remains. Workers describe the job as easy and boring. Some are assigned to cut. Others scan barcodes and pages. There are no complex tasks. The facility has been operating long enough that earlier this year, employees worried it might shut down because they had worked through their supply of books and new shipments were slowing.
The 404 Media investigation, published August 17, 2026, is the first confirmation that Amazon is systematically acquiring physical books for AI training data. Booksellers on marketplaces like Biblio had noticed a spike in large orders over the past year: hundreds or thousands of books at a time, seemingly random titles, from buyers who were not price sensitive. The buyers did not care about condition, edition, or rarity. They wanted content.
---
The ISBN Index
The orders never included books without ISBNs, the unique serial numbers assigned to published works. Booksellers concluded the buyer was working through a list. Not browsing, not selecting, but indexing. The goal was to scan every published book, systematically, one ISBN at a time.
This is a procurement problem, not a literary one. A printed book contains text that is not available on the open internet. It is organized into chapters and sections. It has been edited by humans. And if it was printed before 2022, it is guaranteed to contain no AI-generated text. That last property is now valuable in a way it was not three years ago.
Large language models trained on internet text are increasingly training on text that was itself generated by earlier models. This recursive contamination, called model collapse, degrades model quality over time. The output becomes less coherent, less factual, less useful. The remedy is to find training data that predates the contamination. Printed books, sitting on library shelves and in booksellers' warehouses, are the largest reservoir of uncontaminated human-written text that has not already been scraped.
Amazon is not the only company acquiring books this way. A lawsuit from book authors against Anthropic revealed an internal program called Project Panama. The goal was identical: buy books from commercial marketplaces, cut the spines, scan the pages. The physical book was destroyed in the process. Anthropic argued, and a judge agreed, that this was fair use.
---
The Destruction Incentive
The fair use ruling is where the structural argument lives. The judge ruled that scanning a book for AI training data did not violate copyright in part because the original physical copy was destroyed. Destroying the original meant the copy was not duplicated and resold. The buyer exercised their right to convert physical media to digital format, and by destroying the source, they eliminated the possibility of competing with the publisher by reselling the scanned copy.
This creates a legal incentive that runs in one direction. Keeping the book after scanning it weakens the fair use defense. Destroying it strengthens it. The optimal legal strategy for any company scanning books for training data is therefore: buy, scan, destroy. The destruction is not incidental. It is structurally necessary.
For rare books, the implications are specific. A book with few copies in circulation loses one copy every time it passes through a facility like VGT3. The bookseller who worked with 404 Media made the distinction between types of value. There is monetary value, historical value, intellectual value, sentimental value. The AI company purchasing the book recognizes only one: the text as a sequence of tokens. The physical artifact, the scarcity, the provenance are irrelevant to the buyer and lost in the scanning.
"They just want the content as a bunch of words strung together," the bookseller said.
---
The Purity Premium
Something unusual is happening in the market for information. The internet, which was supposed to make all human knowledge freely available, has become contaminated. AI-generated text now saturates search results, product reviews, news summaries, and social media. The volume of machine-generated content is growing faster than the volume of human-generated content. For AI companies, this creates a paradox: the models produce the pollution that makes their own future training data worse.
The solution is to go backward. Physical books printed before the contamination era are provably clean. They cannot contain AI-generated text because the technology did not exist when they were written. A first-edition paperback from 1987 is now more valuable to an AI company than a freshly written blog post, not because the content is better, but because its provenance is certain.
This is a purity premium. The market is assigning value not to the quality of the text but to the certainty that a human wrote it. A badly written textbook from 1994 is worth more as training data than a well-written essay published yesterday, because yesterday's essay might have been generated or edited by an AI, and there is no way to verify the difference at scale.
The premium inverts the economics of publishing. For five centuries, the value of a printed book declined over time as copies accumulated and newer editions superseded older ones. Now the value curve has reversed. An older book is worth more precisely because it is older. The contamination boundary, roughly 2022, creates a before-and-after that gives pre-contamination books a property no post-contamination text can match: guaranteed human origin.
---
What the Dinosaur Eats
VGT3's mascot is a Tyrannosaurus rex. The choice is probably accidental. Warehouse teams pick symbols that are memorable, not meaningful. But the image works: a creature defined by consumption, holding the thing it is about to consume.
Amazon's Nova family of models competes with frontier models from OpenAI, Google, and Anthropic. Training data is the bottleneck. Every major AI company has already scraped the open internet. The next frontier is text that is not online: corporate documents, private archives, academic papers behind paywalls, and printed books. The company that scans the most books has a training data advantage that compounds. Each scanned book is a marginal improvement to model quality that competitors cannot replicate without scanning the same book.
The books cannot be re-bought after they are destroyed. A rare title with three known copies becomes a title with two. The scanning extracts the text but discards everything else: the typesetting, the marginalia, the physical evidence of how people interacted with the object. The book's text enters the model's weights, is blended with billions of other tokens, and its individual contribution becomes untraceable. The words survive in a form their authors would not recognize.
The bookseller who placed the AirTag watched the tracking dot move across the country, from California to Wisconsin to Colorado to Las Vegas. They knew where the book was going. They did not know what it would become. They knew only that it would not come back.


