When e-books first entered our lives, people said, “Printed books will disappear.”
But few could have imagined that, years later, millions of books would be purchased not for people to read, but to train artificial intelligence models.
Today, we are facing a far more unusual debate:
Is AI really consuming books?
Or is it simply transforming them into a different format?
Why Are Physical Books Becoming AI Training Data?
Developing large language models requires enormous amounts of text data.
Websites, academic research, news content, licensed datasets, and companies’ own data can all form part of this process. Books, however, are particularly valuable to AI companies because they are well structured, have usually been through an editorial process, and often provide in-depth knowledge in specialist fields.
This demand for data appears to be pushing the AI industry beyond digital sources and into the physical book market.
Recent court documents and unusual bulk requests received by booksellers suggest that a surprisingly physical supply chain may exist behind AI training data.
At one end of this chain are major AI companies. At the other are book distributors, second-hand booksellers, and sometimes dealers specialising in rare books.
How Did Anthropic Digitise the Books?
At the centre of the debate is Anthropic, the company behind the Claude AI model.
According to court documents from the US case Bartz v. Anthropic, the company purchased millions of printed books, mostly second-hand copies.
The bindings were removed, the pages were cut to make them suitable for high-speed scanners, and each book was converted into a searchable PDF file. The physical copies were then discarded after scanning.
This method is known as “destructive scanning” because the physical integrity of the book is lost during the process.
The resulting digital files were added to a central research library that Anthropic intended to retain for the long term. The company later selected different groups of books from this collection to train its AI models.
However, the case did not only concern the scanning of physical books.
Court records also revealed that Anthropic had downloaded more than seven million book files from pirated digital libraries such as LibGen and PiLiMi.
Which Uses Did the Court Consider Fair?
In June 2025, Judge William Alsup evaluated three different types of use separately.
The court considered the use of the books to train AI models to be transformative and ruled that it qualified as fair use under the specific circumstances of the case.
Converting legally purchased physical books into digital copies was also considered fair use, but for a separate reason. The court noted that each physical book had been replaced with a single internal digital copy, the number of copies had not increased, and the files had not been distributed outside the company.
However, storing millions of books downloaded from pirated sources to create a permanent digital library was not considered fair use.
This distinction matters.
The ruling does not create a general or international permission allowing an AI company to use any book it purchases in any way it chooses. It applied to specific works, specific uses, and US copyright law.
The remaining part of the case concerning the pirated books was later resolved through a settlement. In July 2026, the court gave final approval to a $1.5 billion settlement covering more than 482,000 works.
The settlement did not overturn the fair-use ruling concerning legally purchased books used for AI training. Instead, it resolved the claims related to copies obtained from pirated sources.
Unusual Requests Reaching Second-Hand Booksellers
The debate is not limited to the Anthropic case.
In 2026, Pieter de Vries, a seller of rare and antiquarian books in Haarlem in the Netherlands, received a highly unusual email.
The sender said they were acting on behalf of a company called 2077AI. Attached to the email was a list containing the ISBNs of 3,001 books. Many of them had been published by academic publishers in 2020 and 2021. Their subjects ranged from business and engineering to education and medicine.
De Vries initially believed the message might be a scam or spam. It later emerged that other Dutch booksellers had received similar requests.
The email reviewed by Fortune did not state how the books would be used. 2077AI also did not respond to questions about whether they were being purchased for AI training.
It is therefore impossible to say with certainty that the books on the list were scanned and destroyed.
However, similar bulk requests received by second-hand booksellers in Germany and Switzerland, involving books from highly specific and unrelated fields, have raised the same suspicions.
It is important to separate the facts from the concerns here.
Court documents confirm that AI companies have purchased and scanned physical books in large quantities. However, it has not been proven that every unusual book order is connected to AI training or that every purchased book is destroyed.
Anthropic has also said that it obtained its books through ordinary commercial channels and that its data collection programmes do not purchase and destroy rare or antique books.
Does Buying a Book Mean Buying All Its Rights?
The legal and ethical sides of this debate may provide different answers to the same question.
When you buy a physical book, you own that particular copy. You can resell it, donate it, write notes in it, or physically destroy it.
However, owning the physical copy does not mean that you also own the copyright to the work.
What makes the Anthropic ruling noteworthy is that the court did not treat ownership of the physical copy and ownership of the copyright as the same thing. Instead, it found the specific use to be transformative and limited.
The ethical question is broader than the legal one.
An action being lawful does not necessarily mean that society will find it culturally acceptable.
An ordinary edition with hundreds of thousands of existing copies may not have the same cultural value as a rare or out-of-print book, or one containing notes left by previous owners.
For this reason, the debate cannot be resolved by asking only whether the book was legally purchased.
We must also consider which book was purchased, how many copies still exist, whether that physical copy has a unique value, and how the digital file will be used.
Does a Digital Copy Preserve Everything About a Book?
A digital copy can preserve the text of a book.
But a book is not always made up solely of the words inside it.
The type of edition, the paper, the cover design, the binding, notes from previous owners, signatures, and the book’s history can all form part of its meaning as a physical object.
This distinction becomes particularly important for rare and historical works.
A high-quality scan can make a book’s content more accessible and help preserve its text for future generations. However, it cannot fully preserve the cultural context of the physical object.
The argument that “nothing is lost as long as the digital copy survives” may therefore not apply equally to every book.
Why Does This Debate Matter for Ireland?
A US court ruling does not automatically mean that the same practice would be considered lawful in Ireland or elsewhere in the European Union.
The European Union’s copyright framework permits text and data mining of lawfully accessible works under certain conditions. However, rights holders may expressly reserve their works from being used for this purpose through appropriate means. The Directive on Copyright in the Digital Single Market contains different rules for scientific research, commercial text and data mining, and the preservation of cultural heritage.
The EU AI Act also requires providers of general-purpose AI models to establish a policy for complying with European Union copyright law and publish a sufficiently detailed summary of the content used to train their models. According to the European Commission, these obligations are intended to make model development and training more transparent.
For Ireland, this issue does not concern AI companies alone.
Authors, publishers, researchers, libraries, archives, and second-hand booksellers may also become part of this emerging data economy.
One practical example of Ireland’s approach to preservation is the legal deposit system. The National Library of Ireland collects publications issued in Ireland so they can become part of the national collection, remain available to researchers, and be preserved for the future.
This approach reminds us of something important:
Societies preserve books not only to access the information inside them, but also because they form part of our published cultural heritage.
Could the Role of Libraries and Booksellers Change?
As AI companies continue to seek high-quality and specialised data, libraries and second-hand booksellers may face new questions.
Can a bookseller be held responsible for how a book is used after it has been sold?
Could rare books require different sales or verification processes?
Will libraries remain institutions that simply preserve knowledge, or could they take a more active role in creating licensed and responsibly sourced datasets for AI training?
Perhaps the real need in the future will not be to collect books through hidden supply chains, but to develop more transparent licensing models involving authors, publishers, libraries, and technology companies.
Such an approach could give AI models access to high-quality knowledge while protecting the rights of creators and the cultural value of physical collections.
The Debate Is Bigger Than Books
Perhaps artificial intelligence is not literally consuming books.
But when we treat books solely as text data to be extracted, we risk losing their physical history, cultural context, and the human stories they carry.
The real question, therefore, is not simply whether physical books are destroyed after being scanned.
It is how we define the value of a book in the age of AI.
If a book has been legally purchased, is it acceptable to digitise it, use it to train AI, and then destroy the physical copy?
Or is preserving the knowledge enough, regardless of what happens to the physical book?


