The AI Rare Book Shredding Panic

The AI Rare Book Shredding Panic (dispatch)

Our read

The panic over AI companies shredding rare books is peak midwit theater. If a book is so rare that it only exists in three moldy basements, digitizing its contents into a neural network that can actually use it is a massive net positive for human knowledge, even if the paper gets recycled.

Published 2026-07-29

Download card
+4

What happened

AI companies are reportedly sourcing rare, out-of-print, and physical books to scan and shred for training data, sparking outrage among archivists and collectors.

The brief

The outrage isn't about saving history; it is about hoarding it. Archivists would rather let a text rot in a locked vault than let an LLM make its contents useful to the public.

The sides

  • The Preservationists

    Physical books are sacred artifacts that must not be destroyed for machine training.

  • The Accelerationists

    Digitizing rare knowledge into a permanent model weights file is the ultimate form of preservation.

Why now

The physical destruction of rare books by AI labs has hit the timeline, triggering a massive wave of outrage from bibliophiles and tech skeptics who view the practice as a literal digital book burning.

Questions

Are AI companies actually shredding rare books to train models?

Yes, some AI labs buy physical out-of-print books, cut the spines off to run them through high-speed scanners, and recycle the paper. This destructive scanning is the fastest way to digitize rare text for training data, turning paper-bound knowledge into machine-readable tokens.

Why do AI labs destroy the books instead of just reading them?

Destructive scanning is a standard industry shortcut for high-volume digitization. Flatbed scanning of bound books is too slow and expensive, so labs slice the bindings to feed pages through automatic document feeders, trading the physical artifact for a perfect digital copy.

Is AI book shredding bad for the preservation of human history?

No, it actually preserves the information. Keeping a rare text locked in a moldy basement where nobody can read it is not preservation, it is hoarding. Digitizing the text into a neural network ensures the actual knowledge survives and remains useful long after the physical paper decays.

Who is actually buying and destroying these physical books for AI training?

Third-party data brokers and specialized scanning operations are the ones purchasing physical books in bulk to feed the AI pipeline. These middlemen buy cheap, out-of-print paperbacks and discarded library stock, slice the spines off for high-speed sheet-fed scanners, and sell the clean text files to AI labs. The tech giants themselves rarely touch the physical paper, outsourcing the manual labor to avoid the bad optics of running a literal industrial book shredder.

Why can't AI companies just use existing digital libraries like Google Books?

Copyright lawsuits and licensing restrictions have locked up the vast majority of existing digital archives. Google Books and the Internet Archive are constantly tied up in federal courts, making their databases a legal minefield for frontier model training. Sourcing physical copies of obscure, out-of-print books and scanning them directly allows AI companies to bypass digital rights management and build proprietary datasets with less immediate legal exposure.

What is the strongest argument against this high-speed scanning practice?

The real risk is the permanent loss of unique historical context that cannot be captured by a flat optical scan. Marginalia, physical binding techniques, paper watermarks, and even the chemical composition of ink hold vital historical data for researchers. When a rare or semi-rare volume is destroyed for its text, we lose the physical artifact forever, trading deep historical evidence for a flat string of tokens in a commercial database.

How does this scanning panic compare to historical preservation efforts?

This is simply the industrial-scale evolution of microfilming and early digital archiving, which have always involved destroying books to save their contents. Throughout the 20th century, libraries routinely destroyed newspapers and rare journals after filming them to save physical shelf space. The only difference now is the speed of the destruction and the fact that the end user is a commercial neural network rather than a university researcher.

What happens to the market value of physical books if this trend continues?

The market value of surviving physical books will likely skyrocket as physical paper becomes a premium status symbol. As more text is digitized and digested by AI, physical libraries will transition from utilitarian information repositories to high-end art galleries. Collectors will pay a massive premium for untouched, physical copies of books precisely because they represent a tangible, unhackable link to pre-AI human culture.

Receipts

Related dispatches

All dispatches · Gifnotes