AI companies are anonymously buying and destroying millions of books through middleman services to avoid headlines about AI companies buying and destroying millions of books

AI isn’t doing great on the PR front. Whether it’s the exponentially increasing electricity consumption amidst worsening climate change, the data centers leaking toxic pollutants into neighboring communities, or the years-long parade of business leaders insisting the technology will replace humans in “most, if not all” professional fields, the AI industry has an excellent record of inspiring resentment and contempt.

And, it seems, the AI companies know it. New reporting from 404 Media suggests that, in an effort to shield themselves from the PR fallout of their data harvesting, major AI providers have turned to third-party middleman services to anonymously source the physical books they’ve been stripping for training data at an industrial scale—and destroying in the process.

In an information ecosystem increasingly poisoned by regurgitated and recirculated AI slop, untainted samples of quality written text have, ironically, only become more valuable for the companies making them so hard to find. The AI model arms race requires an ever-widening set of training data, making physical books that predate the proliferation of LLM-generated text a precious source of pristine, human-authored material to imitate—particularly if it’s a book rare enough to be absent from your competitor’s datasets.

Unfortunately, a book is only valuable for an AI company until its last page has been scanned, as demonstrated by a court ruling in 2025 revealing that Anthropic had carved up, de-spined, scanned, and ultimately discarded millions of print books in the process of training its Claude models. Destroying books, it turns out, isn’t just cheaper than maintaining them: The presiding judge also ruled that it’s transformative enough to constitute fair use under Section 107 of the Copyright Act.

Leave a Comment