Tech companies are currently purchasing and destroying millions of physical books to provide training data for artificial intelligence models. Large-scale operations involve buying rare volumes, stripping the spines from the bindings, and running the pages through high-speed scanners. Once digitized, the physical remains of these books are shredded or pulped. Companies argue this practice prevents their models from consuming low-quality internet content and helps avoid model collapse, a phenomenon where AI systems degrade after being trained on machine-generated output.

Legal proceedings have brought these internal efforts to light. In a notable case involving Anthropic, a federal judge ruled that the destructive scanning of copyrighted books qualifies as fair use. The ruling hinged on the fact that the physical copies were destroyed, with the digital version serving as a direct replacement. Because the digital copy was not distributed and the physical original no longer existed, the court found no violation of copyright law. This legal precedent creates a situation where companies receive more protection by destroying physical media than by preserving it.

Industry critics and former officials express significant concern over these methods. The secrecy surrounding projects like the one internally dubbed Project Panama suggests that these firms recognize the negative public reaction to book destruction. While companies defend the practice as a necessary step for high-quality data acquisition, authors and cultural preservationists view it as an attempt to subsume human history for corporate gain. Despite the controversy, some organizations maintain that they only provide bibliographic data and are not involved in the physical destruction or training of models.

This shift in data acquisition highlights the ongoing tension between copyright law and the requirements of modern machine learning. As firms search for clean, pre-2022 datasets, the demand for print materials that are free from AI-generated text has increased. The resulting legal and ethical debates continue to shape how developers gather the information necessary to update their systems while navigating the consequences of destroying physical archives.