Amazon told 404 Media that it purchases books “to improve our products.” The company reportedly declined to disclose how many books it buys, how many facilities scan them, which products use the resulting material, or whether the Las Vegas workflow is specifically used to train AI models.
That distinction matters. The investigation provides operational evidence: a tracked shipment, a destination inside an Amazon facility, a book-processing operation, and worker accounts describing scanning and destruction. It does not amount to a complete public admission that every book processed there becomes training data for a named Amazon model.
The most accurate conclusion is narrower but still significant: 404 Media found evidence of an Amazon book-digitization operation that appears compatible with AI-data collection, at a moment when Amazon is investing in frontier-model development and data services. Reuters separately reported that Amazon was refocusing its AI strategy around a new frontier-model effort.
Booksellers told 404 Media that some bulk orders looked unusual because they were eclectic and appeared relatively insensitive to a book’s resale value, condition, subject, or price. A buyer interested in resale normally needs to make selections based on demand and margin. A buyer interested in training data may instead care about acquiring a broad range of titles and ISBNs.
That pattern has led dealers to suspect that AI companies or data brokers may be trying to work through the published-book catalog rather than selectively purchasing commercially valuable titles. 404 Media previously reported that ISBNdb marketed high-volume acquisition of printed books for AI companies; after the reporting, ISBNdb removed the relevant material and said the webpage had been a test of market interest.
The evidence does not prove that every unusual order is connected to Amazon or AI. It does show why the market is difficult for booksellers to interpret: an apparently random order can make little sense as a retail transaction while making more sense as an attempt to maximize the breadth of a dataset.
Printed books offer something AI developers increasingly value: large bodies of human-written language created before generative-AI text became widespread. Older and out-of-print works can also contain specialized, local, historical, or otherwise difficult-to-find information that has never been made readily available online.
That makes physical books a potential source of text that is both broad and less likely to consist of synthetic material. 404 Media’s reporting on ISBNdb framed old books as valuable partly because they are “free of AI slop,” reflecting concern that future datasets could contain increasing amounts of AI-generated text.
Researchers and developers also discuss “model collapse,” the possibility that repeatedly training models on synthetic output can reduce the quality or diversity of later models. That concern helps explain the interest in older human-created material, but it is not evidence that Amazon’s models have experienced model collapse or that the Las Vegas books were collected for that reason.
The reported Amazon workflow resembles Anthropic’s Project Panama. In that operation, Anthropic purchased physical books, removed their bindings, scanned the pages, and discarded the paper copies to create digital material for its AI work.
A June 23, 2025 order in Bartz v. Anthropic drew an important distinction. Judge William Alsup found that using the books to train Claude and its predecessors was “exceedingly transformative” fair use under the facts analyzed by the court. He also treated digitizing lawfully purchased print books as fair use because Anthropic replaced purchased physical copies with corresponding, searchable digital copies for its internal library.
The ruling was not a general finding that destroying a book automatically makes AI training lawful. The acquisition method, the purpose of the use, whether copies were retained or redistributed, and the surrounding facts remained important. The same litigation treated Anthropic’s separate creation of a library from pirated books differently.
That legal distinction is central to the Amazon story. Even if Amazon is buying books lawfully, questions about the scale of the operation, the use of the scans, contracts with authors or publishers, and the preservation of rare physical works remain separate from the narrow fair-use analysis in Bartz.
Booksellers’ objections are not limited to lost sales or copyright compensation. Rare and out-of-print books can have historical, intellectual, cultural, and sentimental value. When a scarce copy is cut apart and recycled, the text may survive as a scan, but the physical artifact is gone.
That creates a preservation concern that a copyright ruling does not fully answer. A process can be legally defensible under one set of facts and still alarm collectors, booksellers, libraries, and readers who see physical books as more than containers of text.
The controversy also highlights an information asymmetry. AI companies may understand the dataset value of obscure books long before individual sellers realize that a seemingly odd bulk order is part of a broader acquisition strategy. Amazon’s public statement that it buys books to improve its products leaves open the practical questions that matter most to the book trade: how extensive the buying is, what is being scanned, and what happens to the resulting files.
The alleged book operation fits into Amazon’s broader effort to compete in advanced AI. Reuters reported in July 2026 that Amazon was winding down many in-house models while concentrating resources on a new frontier-model effort. AWS also markets Nova Forge, a service that lets customers combine proprietary data with Amazon-curated training data while developing customized frontier models.
Those facts establish a strategic reason for Amazon to value high-quality text, but they do not independently prove that the books tracked by 404 Media trained a particular model. The strongest evidence remains the investigation’s direct reporting about the shipment and Las Vegas facility.
The lasting question is whether the physical book market is becoming an invisible supply chain for AI. 404 Media’s reporting suggests that, at least in one documented case, books moved from a bookseller to an Amazon processing operation where they were scanned and destroyed. The full scale and downstream use remain undisclosed.