We can’t save all the books. But we can save some of them.
The driving force behind this destruction is what researchers call “model collapse.” When AI systems train on text generated by other AI systems, quality degrades with each generation, producing increasingly incoherent results. The internet has become so saturated with synthetic content that companies now actively seek pre-2022 printed books as the last uncontaminated reservoir of human-authored knowledge.
ISBNdb, a company that sources printed books for AI training data, advertises on its website that “the world’s best AI training data is sitting on a shelf.” The company describes printed books as “curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate.” Pre-2022 publications are “structurally guaranteed to be free of this contamination,” the company states, referencing both AI-generated text and the growing practice of authors “poisoning” web content to sabotage scraping pipelines.
The industrial process of destruction
The mechanics of this operation are precise and troubling. Standard pallets of books hold 800 to 1,200 volumes. Buyers scale from pilot orders to 10,000 or more books per batch. Destructive scanners process 80 to 120 pages per minute after hydraulic cutting machines slice off book spines. The original books are then pulped and recycled.
According to court documents from the Anthropic case, the company’s “Project Panama” spent tens of millions of dollars on this exact pipeline, contracting with Datamation for scanning services. Booksellers have identified these bulk buyers through telltale signs: abnormal volume, subject-agnostic orders spanning history, botany, regional law, and German economics simultaneously, and total indifference to pricing. As one bookseller told 404 Media, “It’s not just the quantity, but the weirdness of the orders.”
Anthropic’s Tom Harvey, who previously helped create Google Books, oversaw the operation. One bookseller who has sold hundreds of books to suspected AI buyers expressed mixed feelings: “It benefits me financially… On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”
The vanishing heritage
The most troubling aspect of this practice is what it means for rare and out-of-print books. Foreign-language volumes, low-circulation academic works, and antique texts with limited surviving copies face potential extinction. When an AI company purchases the last known physical copy of a rare book, scans it, and destroys the original, that volume ceases to exist in the physical world. It becomes data locked inside corporate servers — inaccessible to future scholars, collectors, or the public.
ISBNdb acknowledges the “optics problem” on its website, noting that “‘AI company destroys two million books’ is not a headline that generates sympathy.” The company promises strict nondisclosure agreements on every engagement, stating that “your identity, strategy, and acquisition targets are never disclosed.” This secrecy means booksellers can only speculate about who is buying their inventory and for what purpose.
A legal framework that encourages destruction
Judge Alsup’s ruling explicitly validated the buy-scan-destroy pipeline, writing that “every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other.” Because the digital copy was never shown, shared, or sold outside the company, the judge found this “clearly transformative” and therefore protected by fair use.
This reasoning effectively creates a legal incentive for destruction. If a company keeps the physical book, it must store it. By destroying the original, the company can argue it “replaced” the physical copy with a digital one. The ruling handed the entire AI industry a template, and companies now explicitly cite it.
Harvard University, partnering with Google and Microsoft, demonstrated that an alternative exists. The collaboration released nearly one million public-domain digitized books in 254 languages — no destruction required. Microsoft’s Burton Davis called starting with public-domain data “prudent,” noting that libraries hold “significant amounts of interesting cultural, historical and language data.” That alternative track makes clear that labs have cleaner options. They simply choose not to take them.
We’re going to come up with a program to save and bind some of these rare books. They don’t have to be destroyed. What I’m thinking of is an “adopt a book” program that will permit people to fund the rebinding of a rare book in leather, after which they will receive that book or be recognized as the sponsor of that book in Castalia’s library. If we can just convince the AI companies to send us the pages instead of pulping them, this should be something doable.
In the meantime, you can help us create books that are going to outlive the AI companies engaging in this destruction.
And be sure to take notice of how this is being caused as a downstream consequence of copyright law…

