The recent court case Bartz v Anthropic PBC reveals a significant development in how artificial intelligence companies procure training data. Internal company documents brought to light a secret initiative titled Project Panama. The goal of this effort was the destructive scanning of books to feed high-quality language data into the Claude model. Anthropic sought text created before 2022 to avoid the impact of generative AI on modern writing, relying on books to provide complex, well-organized narratives and facts.
Rather than navigating the traditional copyright licensing process, the company opted for a logistical strategy involving the purchase of physical books. Vendors were hired to slice the spines of these volumes, scan the pages into digital formats, and then discard the physical copies. This process occurred on a massive scale within warehouses, effectively turning printed cultural artifacts into raw data inputs before sending the remnants to recycling or landfills.
Legal analysis of the situation hinges on the fair use doctrine. A judge ruled that using copyrighted material to train an artificial intelligence system does not constitute copyright infringement. The court compared this act to a human reading a book to learn how to write. Because the digital copies were not sold or shared publicly, the court found the transformation of the text into a machine-readable format to be legally acceptable.
This case raises questions regarding the lack of regulation concerning American cultural heritage and the physical treatment of books. While the legal system views these actions as a matter of data acquisition, the practice highlights a shift where large-scale destruction of physical media is treated as a practical alternative to working with authors. As companies continue to search for human-generated text to improve model performance, the status of physical archives remains precarious.
The findings from this litigation suggest a tension between the value of high-quality, human-created literature and the technological demand for massive datasets. As the industry advances, the case serves as a point of reflection on how we preserve our textual history. The reliance on human-generated text to power these systems underscores the continued relevance of the very books that are being destroyed in the process of their ingestion.

