AI companies, led by Anthropic, are buying millions of physical books — including rare and antiquarian titles — running them through destructive spine cutting scanners, and pulping the originals to create clean traini...

Create a landscape editorial hero image for this Studio Global article: What are the key details about AI companies purchasing millions of physical books for destructive scanning to train their models, including. Article summary: I'll research this multi-part question systematically. Topic tags: general, general web, user generated, news. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clickbait thumbnails, icons, and tiny thumbnail layouts. Make it useful as an illustrative visual, not as factual evidence.
Anthropic, OpenAI, and at least one unnamed Chinese AI firm have been quietly purchasing millions of physical books — including rare, out-of-print, and antiquarian titles — running them through high-speed destructive scanners that cut off spines and shred the pages, then sending the remains to pulp. The practice, exposed through court filings, a 404 Media investigation, and the experience of a Dutch bookseller named Pieter de Vries, has sparked a debate about cultural heritage, copyright, and the lengths AI companies will go to for clean training data.
In mid-2026, Dutch antiquarian bookseller Pieter de Vries, who runs De Vries & De Vries in Haarlem, received an email from a woman named Natalia at a Singapore-registered company called 2077AI. The email requested a “fairly large order” of books and included a spreadsheet listing more than 3,000 ISBNs of obscure, pre-2022 titles . De Vries initially dismissed it as spam or phishing
. He later discovered that hundreds of rare-book dealers globally had received similar approaches from companies buying up second-hand books for AI training — often paying 3–5 times market price for rare pre-19th-century works
. The books listed in the spreadsheet were largely published between 2020 and 2021, from academic publishers including Elsevier, Wiley, Routledge, and Oxford University Press
. Topics ranged from business and education to engineering and medicine
.
Court documents unsealed in a federal copyright lawsuit in January 2026 revealed an operation inside Anthropic codenamed "Project Panama" that ran from early 2024 . An internal planning document stated the project’s goal: "Project Panama is our effort to destructively scan all the books in the world"
. The same document added, according to the Boston Globe: "We don't want it to be known that we are working on this"
.
The operation aimed to destructively scan up to 2 million books in roughly six months, hiring a vendor to perform the work . Anthropic spent tens of millions of dollars acquiring millions of physical books, slicing off their spines with a hydraulic cutting machine, running pages through high-speed scanners, and sending the remains to recycling
. The project was led by Tom Harvey, a former Google Books engineer
. Vendor proposals cited in court filings indicate the company sought scanning capacity for 500,000 to 2 million volumes
. The precise final number remains redacted, but filings describe it in the millions
.
Fair-use ruling: US District Judge William Alsup ruled that Anthropic's destructive scanning of legally purchased physical books constituted fair use under copyright law. In his June 2025 order, Alsup described AI training as "exceedingly transformative" and compared the process to "conserving space" through format conversion . The ruling hinged on three facts: Anthropic had bought the books legally, destroyed each copy after scanning, and kept the digital files internal rather than distributing them
.
$1.5 billion settlement: On July 20, 2026, a federal judge approved a $1.5 billion copyright settlement covering pirated books that Anthropic had acquired and stored without authorization — a separate track from the physical-book scanning program . The case revealed that before turning to legal purchases, Anthropic co-founder Ben Mann personally downloaded millions of pirated books
. Judge Alsup drew a clear line: legally purchased books scanned for internal training = fair use; pirated digital copies = liability
. The settlement covered over 91% of eligible authors, according to Anthropic
.
ISBNdb, a company that claims to operate the world's largest book database (over 111 million cataloged titles), now brokers bulk physical-book acquisitions for AI companies . ISBNdb's pitch to AI labs is blunt: "The world's best AI training data is sitting on a shelf" — and its service arranges orders ranging from 1,000 to 1 million books per engagement, under strict non-disclosure agreements
.
The company explicitly markets pre-2022 books as ideal because they are free of AI-generated text . Books, it says, are "dense, edited, authoritative"
. ISBNdb acknowledges that "the optics problem is real" — buyers typically refuse to disclose which lab they represent, and the supply chain is designed to obscure the end customer
. A Change.org petition has been launched demanding ISBNdb stop brokering rare and out-of-print titles for destruction
.
Dutch antiquarian booksellers report that AI companies are acquiring and destroying pre-19th-century rare books — including 18th-century Dutch legal compendia and 17th-century Latin theological manuscripts — at 3–5 times market price . Dealers fear that unique surviving copies of obscure scholarly and regional works are being permanently lost. Unlike mass-market paperbacks, many of these titles exist in only a handful of copies worldwide
.
The process is industrial: books arrive by the pallet, are fed into spine-cutting scanners that can process thousands of volumes per day, and the physical remains are sent to pulp — with no library or archive receiving a copy . Because the destructive scanning happens in secret and the original copies are destroyed, there is no public record of what has been lost, making it impossible to assess the full cultural toll
. Dealers and archivists argue that this lack of transparency means the damage is invisible and irrecoverable
.
Studio Global AI
Use this topic as a starting point for a fresh source-backed answer, then compare citations before you share it.
AI companies, led by Anthropic, are buying millions of physical books — including rare and antiquarian titles — running them through destructive spine cutting scanners, and pulping the originals to create clean traini...