The Copyright Loophole That Turns Book Destruction Into AI Training Strategy
2026-07-29Anthropic bought millions of physical books legally. Then they fed them into a hydraulic-powered cutting machine. Pages removed, spines detached, remnants recycled. The whole operation designed to create training data while avoiding copyright liability. And it worked.
The legal cover came from a June 2025 ruling. In Bartz v. Anthropic, U.S. District Judge William Alsup of the Northern District of California decided that digitising legally purchased print books and using the digital copies to train large language models was fair use. The logic: physical destruction means only one copy exists at any moment. Buy a book, scan it, destroy it, rinse, repeat. The judge noted that the process replaced the print copies with digital versions without adding new copies or redistributing them. One-for-one. Clean.
But here's where it gets perverse. Destruction isn't incidental to the strategy. It's essential to it. If a company kept the book and the digital scan, suddenly there are two copies. That's legally dicier. Anthropic used industrial-grade imaging equipment after the cutting machines did their work. Preservation became legally inconvenient. The law didn't explicitly demand book shredding—it just made shredding the rational choice.
Enter ISBNdb. A company that maintains a large book database, now brokers high-volume acquisitions for AI clients. They help bulk-buy anywhere between 1,000 to one million books per order. Used booksellers started noticing odd patterns. A dealer in Haarlem received a request for nearly 3,000 titles from a Singapore-based firm, while a German seller noted overnight orders from the Canadian company Zoom Books for unrelated academic works.
The archived marketing is almost cynical. ISBNdb boasted about securing books "scattered across library shelves, used bookstores, and out-of-print catalogs," and featured a promise of a "strict NDA on every engagement," ensuring that "your identity, strategy, and acquisition targets are never disclosed." Why the secrecy? The blog post noted that "'AI company destroys two million books' is not a headline that generates sympathy." They knew exactly what they were describing.
Pre-2022 books are the real prize. Physical books that predate the proliferation of LLM-generated text are a precious source of pristine, human-authored material—particularly if it's rare enough to be absent from your competitor's datasets. That's where the economic logic gets tight.
The technical distinction matters: Anthropic's program and ISBNdb's offer remain separate source chains, and neither ISBNdb's marketing nor public reports identify an AI buyer behind a completed order or show that such an order ended in destructive scanning. But the absence of a documented corpse doesn't settle the underlying question.
Judge Alsup's fair use finding, though specific to Anthropic's case, is already shaping how the wider industry approaches sourcing training data from print books. The incentive structure is now locked in. Buy, scan, destroy. That's the efficient path. Anthropic later settled related claims involving pirated digital books for $1.5 billion. The court distinguished between lawfully purchased print books—fair use—and pirated digital collections, which were not. One gets destroyed with impunity. The other gets a settlement.
We don't yet have a named list of specific rare books that vanished. No smoking-gun catalogue of irreplaceable volumes ground to pulp. But the system now rewards what the law didn't explicitly demand. Once books became scalable training data, preservation became economically irrational. Copyright doctrine tilted the balance. Humanity's textual heritage shreds quietly, legally, efficiently.
Source & further reading:
- Microsoft Says MDASH Beats Claude Mythos and GPT-5.6 Sol in Cybersecurity Test — Decrypt
- SpaceX is a battleground Solana must win — CoinDesk
- Live updates: Bitcoin clears $64,000 in Asia hours ahead of Fed decision — CoinDesk
- Company behind AI trade that caused $60 million crypto liquidations to cover all losses — CoinDesk
- Citadel bets on a Fed rate hike Wednesday as bitcoin analysts call a hold. Someone will be wrong. — CoinDesk
- The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions — Dallas Express
- AI Companies Are Still Buying Up Old Books by the Pallet – Then Shredding Them — Yahoo News
- AI companies are anonymously buying and destroying millions of books through middleman services to avoid he... — Yahoo News
- AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale, Even If Almost No Copies Remain — Futurism
- Inside Project Panama, Anthropic's Secret Effort To Scan and Shred the World's Books — IBTimes UK
- AI firms are shredding physical books because copyright law is quietly rewarding them — CryptoSlate
Sources
- Microsoft Says MDASH Beats Claude Mythos and GPT-5.6 Sol in Cybersecurity Test
- SpaceX is a battleground Solana must win
- Live updates: Bitcoin clears $64,000 in Asia hours ahead of Fed decision
- Company behind AI trade that caused $60 million crypto liquidations to cover all losses
- Citadel bets on a Fed rate hike Wednesday as bitcoin analysts call a hold. Someone will be wrong.
- The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions
- AI Companies Are Still Buying Up Old Books by the Pallet – Then Shredding Them
- AI companies are anonymously buying and destroying millions of books through middleman services to avoid he...
- AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale, Even If Almost No Copies Remain
- Inside Project Panama, Anthropic's Secret Effort To Scan and Shred the World's Books
- AI firms are shredding physical books because copyright law is quietly rewarding them