top of page

AI Is Eating the Library: When Training Data Starts Destroying Rare Books

For years, the argument around artificial intelligence and books has centred on a fairly simple question: can AI write a novel? This week, publishing is being forced to confront a more uncomfortable one. What happens when the machinery around AI begins consuming the physical history of books themselves?

Quick Shortcuts

Rare books and AI training — Wall Street Journal reporting

AI-facilitated trade publishing — The Bookseller

Transparent AI use in publishing — Nature Computational Science

Rare books are becoming AI training material

Fresh reporting from the Wall Street Journal describes an extraordinary new pressure on the book trade: obscure bulk buyers have been purchasing rare and out-of-print books, with some volumes reportedly being cut apart, scanned and recycled so their contents can be used as AI training data. For books that exist in plentiful copies, digitisation is one thing. For scarce editions, signed works or potentially unique material, destruction is something else entirely.

There is a strange circularity here. Artificial intelligence was built partly by learning from human culture. Now the demand for ever more specialised training material appears capable of creating a commercial incentive to physically dismantle pieces of that culture in order to digitise them.

This is bigger than another copyright argument

Authors have already spent years debating whether technology companies should be permitted to train models on copyrighted books, whether writers should be paid, and what fair use should mean in an age of generative AI. Those questions remain unresolved. But the rare-book story introduces another issue: preservation.

A book is not only a container for words. A particular edition can be an historical object. Marginalia, inscriptions, bindings, printing errors, provenance and even the physical materials can tell researchers something that a clean text file cannot. Once a scarce physical copy has been sliced apart for scanning, some of that information may be gone permanently.

At the same time, publishing is experimenting with declared AI

The timing is fascinating because another experiment is happening at the opposite end of publishing. Harriman House recently announced The Everything Code by Raoul Pal, describing it as a trade-publishing first: a “human-originated, AI-facilitated” book produced with a proprietary system trained exclusively on Pal’s own existing work.

That phrase matters. Instead of hiding AI involvement, the publisher is making the method part of the book’s provenance. Whether readers embrace that approach remains to be seen, but it offers a much healthier principle than pretending every use of AI is identical.

Where I draw the line as an author

For me, authorship remains about creative responsibility. I write my fiction. The characters, worlds, decisions, rewrites and final creative judgement are mine. I can also use artificial intelligence around the business of being an author — for research support, visualisation, advertising, online content, organisation and other practical tasks. Those activities do not suddenly make the machine the author of my novels.

The distinction becomes even more important as the technology spreads. If we simply label everything “AI” we lose the ability to discuss what actually happened. A spelling tool, a research assistant, an image generator, a marketing workflow and a system generating fifty thousand words of a novel are not the same creative act.

Publishing needs provenance in both directions

Much of the recent debate has focused on proving that a manuscript came from a human writer. But provenance should also apply to the material AI companies consume. Where did the training material come from? Was it licensed? Was a physical artefact destroyed to obtain it? Does the owner understand what they are selling into?

Nature Computational Science argued this month that transparency, accountability and human oversight should remain central as AI becomes embedded in publishing. That principle travels well beyond academic journals. If publishers expect authors to explain meaningful AI use, technology companies should face equally serious questions about the provenance of the books and cultural material used to build their systems.

The opportunity is still real

None of this requires writers to become anti-technology. AI can remove tedious work, open creative possibilities and help independent authors compete in areas that once required teams of people. I use it because some of those capabilities are genuinely useful.

But useful technology does not need a free pass. The more powerful these systems become, the more important it is to know where their knowledge came from and what was sacrificed to obtain it. Digitising human culture can preserve knowledge. Destroying scarce pieces of that culture simply to make another dataset richer is a very different proposition.

Today’s question

If a rare or historically important book is destroyed so its contents can train an AI system, has the technology preserved human knowledge — or consumed part of it? And should AI companies have a responsibility to prove the provenance of their training material just as authors are increasingly expected to prove the provenance of their writing?

About Rob Frankson

Rob Frankson is a science-fiction author and creator of the Near Galaxy Saga. Through 121 Minutes he writes about storytelling, publishing, creativity and the changing relationship between authors and artificial intelligence.

Keep up with AI & The Author

Want these briefings direct to your inbox? Join the 121 Minutes Universe for updates, articles and writing resources.

Explore more original science fiction, novels, short stories and author resources at 121 Minutes.

AI & Editorial Transparency

AI & The Author is edited and published by Rob Frankson. Artificial intelligence is used to assist with news research, initial drafting, content organisation and supporting imagery. All articles are reviewed and, where necessary, edited by Rob Frankson before publication. The opinions, editorial position and final decision to publish remain the author's.

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page