Current Affairs Security

The Copyright Loophole That Turns Book Destruction Into AI Training Strategy

2026-07-29

Anthropic bought millions of physical books legally. Then they fed them into a hydraulic-powered cutting machine. Pages removed, spines detached, remnants recycled. The whole operation designed to create training data while avoiding copyright liability. And it worked.

The legal cover came from a June 2025 ruling. In Bartz v. Anthropic, U.S. District Judge William Alsup of the Northern District of California decided that digitising legally purchased print books and using the digital copies to train large language models was fair use. The logic: physical destruction means only one copy exists at any moment. Buy a book, scan it, destroy it, rinse, repeat. The judge noted that the process replaced the print copies with digital versions without adding new copies or redistributing them. One-for-one. Clean.

But here's where it gets perverse. Destruction isn't incidental to the strategy. It's essential to it. If a company kept the book and the digital scan, suddenly there are two copies. That's legally dicier. Anthropic used industrial-grade imaging equipment after the cutting machines did their work. Preservation became legally inconvenient. The law didn't explicitly demand book shredding—it just made shredding the rational choice.

Enter ISBNdb. A company that maintains a large book database, now brokers high-volume acquisitions for AI clients. They help bulk-buy anywhere between 1,000 to one million books per order. Used booksellers started noticing odd patterns. A dealer in Haarlem received a request for nearly 3,000 titles from a Singapore-based firm, while a German seller noted overnight orders from the Canadian company Zoom Books for unrelated academic works.

The archived marketing is almost cynical. ISBNdb boasted about securing books "scattered across library shelves, used bookstores, and out-of-print catalogs," and featured a promise of a "strict NDA on every engagement," ensuring that "your identity, strategy, and acquisition targets are never disclosed." Why the secrecy? The blog post noted that "'AI company destroys two million books' is not a headline that generates sympathy." They knew exactly what they were describing.

Pre-2022 books are the real prize. Physical books that predate the proliferation of LLM-generated text are a precious source of pristine, human-authored material—particularly if it's rare enough to be absent from your competitor's datasets. That's where the economic logic gets tight.

The technical distinction matters: Anthropic's program and ISBNdb's offer remain separate source chains, and neither ISBNdb's marketing nor public reports identify an AI buyer behind a completed order or show that such an order ended in destructive scanning. But the absence of a documented corpse doesn't settle the underlying question.

Judge Alsup's fair use finding, though specific to Anthropic's case, is already shaping how the wider industry approaches sourcing training data from print books. The incentive structure is now locked in. Buy, scan, destroy. That's the efficient path. Anthropic later settled related claims involving pirated digital books for $1.5 billion. The court distinguished between lawfully purchased print books—fair use—and pirated digital collections, which were not. One gets destroyed with impunity. The other gets a settlement.

We don't yet have a named list of specific rare books that vanished. No smoking-gun catalogue of irreplaceable volumes ground to pulp. But the system now rewards what the law didn't explicitly demand. Once books became scalable training data, preservation became economically irrational. Copyright doctrine tilted the balance. Humanity's textual heritage shreds quietly, legally, efficiently.


Source & further reading:

Sources