Who Owns the Books That Train AI? Copyright’s Next Battle for Authors
- Rob Frankson
- Aug 26
- 4 min read
The argument over artificial intelligence and books has moved beyond whether AI can write. Today, governments, publishers and authors are confronting the harder question: who gets to decide which books can be used to train the machines in the first place?
Quick Shortcuts
Copyright is becoming the battleground for AI and books
On 26 August, Singapore launched a public consultation asking a question that sits at the centre of the global AI argument: can copyrighted works be used to train artificial intelligence, and what protections should creators have when that happens? It is another sign that the debate is moving from theoretical arguments into actual rules.
For writers, this matters because books are unusually valuable training material. They are long-form, edited, structured expressions of human thought. That makes them useful to AI developers and makes the question of permission, payment and provenance impossible to ignore.
The physical book has entered the argument
The controversy has become stranger because AI training is no longer only about digital files. Recent reporting has described large-scale acquisition and destructive scanning of physical books for training data. Civil-society groups have even urged the US Federal Trade Commission to investigate whether the practice raises competition concerns.
That creates an uncomfortable image for anyone who loves books: a physical volume being cut apart so its pages can be scanned, digitised and absorbed into a machine-learning dataset. Digitising a common modern book is one thing. Destroying a scarce edition, annotated copy or culturally significant object is another.
A book is not always just the words printed on its pages. Edition, binding, marginalia, inscriptions and provenance can themselves carry information. Once the object is destroyed, some of that history may disappear even if every printed sentence survives as data.
Writers are being asked for transparency. AI companies should face the same test
Publishing is increasingly demanding clarity from authors about AI. Oxford University Press requires responsible and transparent use and makes authors accountable for the integrity and originality of their work. Amazon KDP distinguishes between AI-generated and AI-assisted content: generated material must be disclosed, while assisted material does not require the same declaration.
Those distinctions are important. They recognise something that often gets lost in the louder argument: using AI somewhere in a creative career is not automatically the same as handing authorship to a machine.
I write my books. I also use AI around the creative business for visualisation, advertising, online content, research support and organisation. For me, that boundary matters. Technology can support an author without becoming the author.
But transparency cannot sensibly operate in only one direction. If writers are expected to explain how AI touched their work, AI companies should be able to explain where the material used to train their systems came from, whether it was licensed, and what happened to the original source.
The real issue is provenance
This is why provenance may become one of the defining ideas of the AI era. For an author, provenance means being able to show how a manuscript developed: notes, drafts, revisions and creative decisions. For an AI system, provenance should mean being able to account for the material that helped shape the model.
Neither side needs a perfect forensic record of every keystroke or every token. But a creative economy built entirely on “trust us” will struggle when valuable intellectual property is involved.
The encouraging part is that policymakers are beginning to ask these questions directly. Singapore’s consultation joins a much wider international effort to find a workable balance between AI innovation and creators’ rights. The UK has already published its own detailed report and impact assessment on copyright and artificial intelligence.
What authors should do now
Writers do not need to become copyright lawyers. But we should pay attention to the terms attached to the tools we use. Do not casually upload unpublished manuscripts into services whose terms allow reuse for training. Keep sensible records of your creative process. Understand the distinction between AI-generated and AI-assisted material on the platforms where you publish. And be clear with readers and publishers about where you personally draw the line.
Most importantly, do not allow the debate to collapse into two camps. “AI everything” and “AI nothing”. The interesting territory is between them: how we use powerful technology while preserving authorship, creative ownership and trust.
Today’s question
If authors are expected to prove the provenance of their writing, should AI companies be required to prove the provenance of the books and creative works used to train their models?
About Rob Frankson
Rob Frankson is a science-fiction author and creator of the Near Galaxy Saga. Through 121 Minutes he writes about storytelling, publishing, technology and the changing creative landscape.
Stay with 121 Minutes
AI & Editorial Transparency
AI & The Author is edited and published by Rob Frankson. Artificial intelligence is used to assist with news research, initial drafting, content organisation and supporting imagery. All articles are reviewed and, where necessary, edited by Rob Frankson before publication. The opinions, editorial position and final decision to publish remain the author's.
.png)


Comments