The race for artificial intelligence supremacy is leading to a controversial practice. Companies are acquiring millions of physical books to feed their data models, often opting to slice the spines off and scan the pages for speed. While this approach provides large sets of high-quality data, it results in the destruction of the original physical copies once they are processed.
Recent court documents have brought these methods to light, showing that major industry players have invested millions in these operations. One vendor even marketed high-speed cutting machines to facilitate this transition from paper to data, encouraging companies to focus on the recycling aspect rather than the destruction of the books.
This trend has drawn sharp criticism from the public and legal experts who point to the moral and cultural loss of such practices. While some companies have faced backlash and moved to pause or pivot away from these destructive models, the underlying demand for massive, reliable training datasets continues to drive interest in these methods.
Nonprofit institutions like the Internet Archive offer a contrast to these industrial practices by using manual, nondestructive scanning methods. However, the speed of manual scanning cannot compete with the massive output of industrial cutting machines, leaving the future of physical book preservation in tension with the rapid development of large language models.

