Evidence is mounting that major tech firms are systematically purchasing rare and antique books to train artificial intelligence models. Recent investigations reveal these companies acquire physical copies of books, hire contractors to remove their spines, and feed the pages into high-speed scanners before shredding the original works. This practice has sparked an outcry from archivists, historians, and authors who view the destruction of rare print copies as a significant cultural loss.
Legal documents and tracking investigations confirm the scale of these operations. One initiative, known as Project Panama, involved plans to scan millions of books. Reporters recently used tracking technology to follow a shipment of rare volumes from a bookseller directly to an Amazon-owned warehouse facility. Inside, the primary function of the staff is to scan books for industrial-scale data ingestion.
AI developers argue this process is a matter of data quality and legal compliance. By purchasing physical copies, they obtain the legal right to digitize and destroy them under the First-Sale Doctrine of the Copyright Act. Furthermore, firms prioritize books printed before 2022 to ensure their training data remains free from AI-generated text, which can degrade model performance. The physical books provide human-authored reasoning and narrative structure that digital-only archives often lack.
While critics decry the loss of physical history, companies maintain that this approach is a necessary step to advance language models. The exact number of destroyed books remains unknown, but reports suggest massive bulk orders of hundreds of thousands of titles are standard. As these tech giants continue to secure high-quality training sets, the intersection of technology and physical preservation remains a point of intense public and legal debate.

