AI Companies Buy and Destroy Millions of Physical Books to Feed Chatbot Training Data
Key Takeaways
- •AI companies are using anonymous intermediaries to purchase physical books at scale, scan them for training data, and discard the originals.
- •Pre-2023 books are especially prized by AI developers because they contain only human-authored text, avoiding contamination from AI-generated content.
- •Booksellers report that weekly sales have risen from roughly 20 books to several hundred due to AI buyers, but some worry that uncommon titles are being permanently lost.
- •A federal judge ruled in Bartz v. Anthropic PBC that destructive scanning of legally purchased books constitutes transformative fair use.
- •Courts have drawn a legal line between scanning purchased books and using pirated copies, as shown by a $1.5 billion settlement requiring Anthropic to compensate authors whose pirated works were used to train Claude.

Artificial intelligence companies are anonymously purchasing physical books in bulk and destroying them after scanning the contents for AI model training, according to a report first published by 404 Media.
Booksellers say demand for obscure and out-of-print titles has surged, raising concerns that rare and uncommon books are permanently disappearing as a result of the practice.
A federal judge has ruled that destructive scanning of legally purchased books can qualify as fair use, even as separate copyright litigation against AI developers continues to work through the courts.
A Cottage Industry for Destructive Scanning
Evoking the classic dystopian novel Fahrenheit 451, some AI companies are not merely reading books—they are destroying them. As developers race to build more powerful AI models and copyright lawsuits pile up, a cottage industry has emerged to supply them with millions of physical books that are stripped apart, scanned into training datasets, and then discarded.
AI companies are reportedly using intermediaries to acquire books at industrial scale while remaining anonymous. Firms specializing in bulk sourcing advertise the ability to locate hundreds of thousands of titles while promising confidentiality for their AI clients, reflecting the sensitivity surrounding the practice.
Critics say the buying spree is driven by AI companies seeking to preserve human-authored knowledge before it becomes diluted by AI-generated text—often referred to as "AI slop". Books published before the rise of generative AI in 2023 are considered especially valuable because they provide high-quality training data written entirely by humans. The push reflects a broader industry challenge sometimes called the "data wall"—as publicly available internet text becomes increasingly saturated with AI-generated output, developers are turning to offline sources of verified human writing to maintain model quality.
Reshaping the Used-Book Market
The surge in demand is already reshaping the used-book market. One unnamed bookseller told 404 Media that weekly sales climbed from roughly 20 books to several hundred after AI buyers entered the market. While the increase has been profitable, the bookseller expressed concern that uncommon and out-of-print books are being permanently lost after they are scanned and destroyed.
"It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell," the bookseller told 404 Media. "I've been well suited for these sales with inventory from overseas and foreign language books. On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped."
The loss of physical copies is particularly acute for books that exist in limited print runs or were never digitized, raising questions about whether AI companies bear any obligation to preserve the artifacts they consume.
Legal Precedent: Fair Use for Destructive Scanning
The practice mirrors Anthropic's "Project Panama", which digitized millions of books through destructive scanning.
Last summer, in the copyright lawsuit Bartz v. Anthropic PBC, a federal judge in San Francisco ruled that scanning legally purchased physical books into digital copies—even when the originals were destroyed—constituted transformative fair use. Federal judges later issued similar fair use rulings in separate copyright cases involving OpenAI and Meta.
However, in a separate case, a federal judge in the same district approved a $1.5 billion copyright settlement requiring Anthropic to pay thousands of authors approximately $3,000 per book after the company used pirated copies of their works to train Claude. The contrasting outcomes underscore a legal distinction that AI companies now navigate carefully: purchasing and scanning a physical book may be protected, but obtaining the same content through piracy is not.
Industry Backlash and Calls for Preservation
In response to growing backlash, AI developers including Elon Musk have spoken out against the practice and urged companies to preserve the books being scanned.
"I've asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning," Musk wrote on X.