Tech giants are turning to a shadow industry of data brokers to fuel their artificial intelligence models, according to recent reports. These companies are allegedly hiring third-party contractors to acquire and digitize millions of copyrighted books, only to destroy the physical copies once the text is ingested into training sets.
The practice bypasses standard licensing agreements by leveraging the “middleman” loophole. By outsourcing the acquisition to smaller firms, tech companies distance themselves from the direct harvesting of intellectual property, creating a buffer against copyright litigation.
Authors and publishers are now sounding the alarm. For writers, this isn’t just about lost royalties; it’s about the ethics of “disposable literature.” Books are being treated as raw industrial fuel—scanned for their linguistic patterns and then discarded like scrap metal.
“They are essentially laundering the data,” said a legal analyst familiar with the ongoing intellectual property debates. “By shifting the scanning process to a subcontractor, the big players claim they didn’t ‘copy’ the books themselves, even though the end result is a model trained on stolen labor.”
The process follows a grim routine: contractors secure thousands of titles from various sources, digitize the pages into high-resolution text files, and then physically incinerate or pulp the books to prevent any possibility of audit or resale. This ensures the physical evidence of the data acquisition vanishes.
Major AI developers have maintained a wall of silence on these specific operations. While they publicly tout “ethical AI” and “fair use” as their guiding principles, the reality on the ground suggests a desperate, high-stakes race to consume as much human-created content as possible before copyright laws catch up.
For the publishing industry, the challenge is proving the source of the training data. Because the books are destroyed and the middlemen operate under strict non-disclosure agreements, tracing a specific model’s output back to a copyrighted source is becoming nearly impossible.
As the industry pivots toward these clandestine methods, the legal definition of “fair use” is being tested in real-time. If the courts decide that transforming a physical book into a digital training set constitutes infringement, the entire foundation of current generative AI models could be at risk.
For now, the machines continue to read, and the books continue to disappear.
