Antique hard copy the best antidote to AI slop
Image: Eugenio Mazzone @eugi1492 via Unsplashed
Silicon Valley tech giants are targeting hundreds of rare book dealers around the world in their drive to scan and pulp mountains of second-hand books to train their large language models.
Anthropic led the charge when it hired a former Google executive to obtain “all the books in the world”. Millions were sourced by a company called ISBNdb, which claims to have “the world’s largest book database”. As the tech news site 404 Media reports, this “makes it easier for AI companies to methodically acquire, scan, and turn printed books into training data while avoiding duplication”.
Internal Anthropic documents that surfaced thanks to a copyright lawsuit brought by authors reveal that the tech company was running a secret programme called Panama Project “to destructively scan all the books in the world”. This equates to chopping off a book’s spine to make it easier to feed pages into an industrial scanner, then pulping them – the quickest and cheapest method of ingesting their content.
The fact that this has become common practice doesn’t mean anyone wants it publicised. “We don’t want it to be known that we are working on this,” Anthropic said in an internal planning document filed during legal proceedings. “The optics problem is real,” ISBNdb says on its website, pointing out that the headline “AI company destroys two million books” is not one likely to generate sympathy.
For this reason, mass book sales to AI companies for scanning and pulping are a secretive business. Strict, legally binding NDAs govern all transactions, and the buyer’s name is kept hidden. “Your identity, strategy and acquisition targets are never disclosed,” ISBNdb’s promotional material promises.
The company is quick to rebuff criticism that the practice amounts to “book burning”, stressing the importance of “responsible physical sourcing”, followed by recycling the discarded pages. “It is the completion of the book’s lifecycle: from tree to knowledge to tree again.”
The reason physical books are in such demand, as ISBNdb points out, is that they “represent curated, peer-reviewed, domain-specific human knowledge structured in a way no web crawl can replicate. Dense, edited, authoritative”.
Those published before 2022 are particularly valuable because they are guaranteed to be free from what’s called “data poisoning”. In an article titled “How Authors Are Fighting Back Against AI Training, And What It Means for Your Data Strategy”, ISBNdb explains writers who objected to their work being scraped began to experiment with ways to sabotage attempts by AI models to ingest it. “Instead of trying to prevent your work from being scanned, you make the work itself cause harm to any model that ingests it, while remaining perfectly readable to any human,” says ISBNdb. The only effective defence is “source discipline: knowing where your data comes from before it enters the pipeline.”
Perhaps even more significant, pre-2022 books are free of AI slop. ISBNdb explains when AI trains on AI, “a documented phenomenon called model collapse occurs: subtle linguistic nuances vanish, systematic errors compound, and outputs converge on repetitive patterns. Each generation trained on synthetic data is slightly worse than the last.” This makes “pre-LLM books uniquely valuable” – a currency that continues to appreciate “as web content quality continues to decline”.
Perhaps the time has come to dust off our old library cards.