Tech companies are buying vast quantities of physical books, only to destroy them after scanning their contents. Large language model providers seek high-quality human text to train their systems. They want to avoid model collapse. This occurs when AI models train on synthetic data generated by other machines, which leads to degraded outputs and loss of accuracy over time.

The Surge in Book Acquisitions

Book dealers first noticed the trend when mysterious buyers began placing large orders for obscure titles. These buyers purchased everything from medieval legal records to 1960s Swedish comedy collections and 2010s Texas civil procedure manuals. Price was not a deterrent for these entities. Sellers often accepted whatever amount was requested for items they previously considered impossible to unload.

404 Media recently traced one such shipment. The books ended up at an Amazon-run scanning facility in Las Vegas. The logistics suggest a systematic effort to strip-mine physical libraries for training material. While these books hold little value to collectors or the average reader, they are essential fuel for LLM developers.

The Destruction Process

Speed is the primary requirement for these scanning operations. To achieve the necessary throughput, workers cut the spines off the books. This allows them to feed the loose pages into high-speed scanners. Once the digital copy is captured, the physical book is discarded or destroyed. The process effectively ends the lifespan of the original print object.

Amazon and Anthropic have both been identified as participants in this practice. Other firms with large-scale model development needs likely follow similar protocols. The activity represents a literal transformation of physical literary culture into raw data. When models run out of new internet content, they pivot back to the physical world for untainted text.

Future Consequences and Industry Implications

Model collapse remains a significant threat for AI firms. If these models train on their own output, their logic becomes circular and flawed. Companies believe using human-authored books provides a necessary buffer against this decay. This strategy keeps them tethered to older, verified information sources.

Still, this practice creates a new tension within the publishing and archival sectors. Libraries and individual dealers face a market where obscure history is suddenly in high demand but destined for the trash. As the need for training data grows, it remains unclear how much of the historical print record will survive the transition into digital model training. Observers note this is an on-the-nose example of technology devouring the culture it relies on for its initial development.