OwnGlobal
Technology

Intent not malicious, it's indifference: Thousands, if not millions, of books sent to AI woodchipper

Intent not malicious, it's indifference: Thousands, if not millions, of books sent to AI woodchipper

This practice disproportionately affects niche, out-of-print, or non-commercial works that lack digital advocacy

Investigative journalist Emanuel Maiberg joins François Picard to discuss how vast quantities of books are being discarded or repurposed to train artificial intelligence systems, often without regard for their cultural or intellectual value. The conversation highlights a growing trend where publishers, libraries, and institutions offload surplus texts, treating them as raw material rather than knowledge assets. This process, described as feeding an „AI woodchipper,”reflects systemic indifference rather than deliberate harm. The scale of book disposal for AI training raises concerns about the loss of diverse perspectives and the erosion of physical archives in the digital age. Maiberg explains that while some digitization efforts aim to preserve content, many books are simply shredded or pulped after scanning, their contents used to train language models without compensation to authors or publishers.

This practice disproportionately affects niche, out-of-print, or non-commercial works that lack digital advocacy. How Do Libraries Decide What Gets Scrapped? Libraries often face space constraints and budget pressures, leading them to prioritize recent or high-demand titles for retention. Older or duplicate copies are frequently deemed expendable, especially when digital versions exist. Maiberg notes that these decisions are rarely transparent, with little public consultation on what constitutes culturally significant material worth saving. What Happens to the Knowledge Inside These Books? Once scanned, the text is broken down into data points for AI training, stripping away context, annotations, and physical form. While the information may persist in model outputs, the original intent, historical nuances, and bibliographic details are often lost.

Picard questions whether this trade-off serves progress or merely convenience, warning that indifference to preservation could impoverish future research. Frequently Asked Questions Why are books being used for AI training instead of being donated? Many institutions cite logistical challenges and lack of demand for physical donations, opting instead for recycling or pulping as a simpler solution, even when better alternatives exist. Are authors compensated when their books train AI systems? In most cases, no—current copyright frameworks do not adequately address the use of published works in AI training data, leaving creators without recourse or remuneration. Could this trend affect historical research? Yes, the loss of physical copies and contextual metadata may hinder scholars’ ability to verify sources or study publishing history, particularly for marginalized voices already underrepresented in digital collections.

Content written by François PICARD for OwnGlobal editorial team, AI-assisted.

Comments (0)