
Last week, the internet raged as social-media posts and news articles accused AI companies of “destroying the world’s books”—including “millions of rare” ones—by chopping off their spines to scan them more easily, and discarding them afterward. One article said the news was “sparking concerns that the last remaining copies of out-of-print texts are being destroyed.” The investor and AI skeptic Michael Burry called the practice of sacrificing rare books “evil incarnate.” Even Elon Musk weighed in, posting that he had asked the engineers training xAI’s models to “preserve any rare books in a library and scan them the hard way.”The destruction of physical books in the process of training AI has been public knowledge for more than a year, but debate over the practice intensified after an article by 404 Media, which described a large number of orders coming to booksellers, suggested that rare books might be included, and named a book-database company called ISBNdb as the possible buyer.Lost in the outrage, however, are several unanswered questions. The 404 Media story said ISBNdb had advertised bulk buying of books for AI companies, but acknowledged that no evidence directly connects the company to the recent wave of purchases that booksellers have described. And although the story calls out rare-book sellers, it’s unclear how rare most of the books being bought actually are, or whether they’re of great value.Despite that, the conversation rapidly intensified into a panic. So what is actually going on? And who is actually responsible?[Read: ]Is This What Comes After AI Slop?Many of the specifics of this situation are obscured by hard-to-trace purchasing histories and the silence of AI companies and other potential buyers, so let’s lay out what we know. We know that at least one major AI company has used physical books to train their models. Documents unsealed earlier this year in a lawsuit against Anthropic reveal that in the spring of 2024, the AI giant bought millions of new and used books and destroyed them in order to scan the pages to train its AI models. Anthropic and other AI companies are likely keen on books published before the launch of ChatGPT, in 2022, which are all but guaranteed to be free of AI-generated text. Although AI-generated text is readable—if often unpleasant—to humans, it can be devastating to AI models: When trained on their own output, they can degrade significantly in quality in a phenomenon known… [TheTopNews] Read More.
1 week ago





