Mediaweek
Vinyl Media

Our Sites

Logo Rolling StoneLogo VarietyLogo PedestrianLogo Refinery29Logo BuzzfeedLogo TastyLogo PopsugarLogo LadbibleLogo SportbibleLogo TimeoutLogo Concrete PlaygroundLogo MediaweekLogo The Music Network

Network Partners

Art NewsBGRBillboardCrunchyrollDeadlineDirtEnthusiast GamingFootwear NewsFunimationGamelancerGold DerbyHypebeastIndieWireKidoodleLife Without AndySheKnowsSourcing JournalSporticoSPYStyleCasterThe Hollywood ReporterToon GogglesTVLineVibe

AI companies accused of buying and destroying books for training data

Booksellers are quietly offloading bulk lots to AI labs, where spines get sliced clean off before the pages get scanned.

By Simran PasrichaPublished Aug 2, 2026
5 min read
MW 030826 2785
TikTok

Another day, another dystopian choice by tech companies. This time, it seems as though AI companies are buying books to scan and train their LLMs (large language models, aka the systems behind tools like ChatGPT) and then destroying them.

Naturally, the internet is furious. Here’s what’s actually going on.

A video of books being sliced by machines is doing the rounds on the internet. One TikToker commented called it “just another form of book burning”.

BookTok creator and author Timna Devolpi went further, saying, “This video is one of the most horrifying things I have ever seen… The written word is all we have.

 

So what’s happening with the books?

The latest flare-up started with reporting from 404 Media, which found that AI companies were quietly buying large numbers of physical books through intermediaries. Booksellers noticed the pattern pretty quickly: big bulk orders, little rhyme or reason, and plenty of ISBNs. Not a lot of signs that anyone was building a wholesome home library.

It seems that this has become an entire business model with ISBNdb, a company that says it has the “world’s largest book database”, now offering bulk book buying for AI labs, including orders of up to one million titles at a time.

mediaweek
Morning Report

The leading media trade publication in Australia.

Get our top stories straight to your inbox daily by signing up to our Newsletter

By providing your information, you agree to our Terms of Use and our Privacy Policy. We use vendors that may also process your information to help provide our services.

Its pitch on their website includes “older, rare and specialist volumes”, with the company saying, “the world’s best AI training data is sitting on a shelf”.

Their sell: “Print books from the pre-LLM era are structurally guaranteed to be free of this contamination.” They also claim that “millions of the most valuable books have never been digitised”, adding: “They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”

ISBNdb has since taken down that marketing page and written a notice on their website claiming the company “has never purchased, scanned, or destroyed a book — for AI training or anything else”.

“We don’t train AI models, and we never have. The page was up to explore demand for a service we never brought to life. We’ve taken it down,” the statement continued.

An updated message.(Image: ISBNdb)

One (separate) seller who has partaken in the process of bulk selling told 404 Media, “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell… On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”

And pulped they are. The books are being used in a process called destructive scanning, which is as brutal as it sounds. The spine is cut off so the pages can be fed into industrial scanners, and once the text has been digitised, the physical book is discarded or recycled.

Court documents from the Anthropic case (more on this below) describe a “hydraulic-powered cutting machine” being used to cut books apart before scanning.

 

Why AI companies want books

This part is less mysterious: they simply need more data. AI models need huge amounts of human writing to train on, and the open internet has increasingly become a sinkhole full of AI slop. So companies have gone hunting for cleaner source material, and books (especially older ones) are appealing because they were written before the internet turned into more of a hellscape.

The books are being purchased because they are useful. They are being destroyed because, from a technical point of view, that is faster and cheaper than keeping the physical copy intact.

It's understood that this practice seems to be widely used among many AI companies and has not been without prior backlash.

 

Anthropic’s court case

Anthropic, the company behind Claude, was sued in 2024 by authors after they accused the company of using books without permission to train its chatbot. Reuters reported in July that a US judge approved the company’s $1.5 billion settlement in that case after the court found that training AI on books is allowed under copyright law.

While this case was still ongoing, court filings revealed that the company launched Project Panama — an effort to buy books in bulk, cut off their spines and scan the pages for model training, with internal documents stating “we do not want it to be known that we are pursuing this project”, per news.com.au.

P.TV understands that Anthropic’s data acquisition programs use a mix of publicly available web data, commercially acquired datasets and data the company generates itself.

When approached for comment on the recent outrage online over the practice, an Anthropic spokesperson asserted to P.TV that they have internal guidelines on the books they purchase and “destroy”.

“None of our data acquisition programs buys and destroys ‘rare’ or ‘antiquarian’ books,” they said.

91 per cent of the more than 480,000 authors and publishers covered by the settlement have claimed their share of the payment.

 

What Australia is saying

Although there are no reports of this process happening in Australia yet, the current laws around copyright and AI in Australia are still grey… at least for now.

Australian copyright law does not have a specific exception for training AI on books, and the usual fair dealing exceptions were not designed for machine learning. That means using copyrighted material to train AI can still require permission from the rights holder, while the government continues to consider whether the law needs updating.

Prime Minister Anthony Albanese recently said in his announcement of the National AI Plan, “Australian writers, musicians, artists, and journalists must retain ownership and control of their work. Our laws will spell that out plain as day.” So I guess we’ll wait and see.

Anthony Albanese

The backlash is landing in a year when book culture is already under pressure. In the US, the American Library Association said 5,668 books were banned in libraries in 2025, the highest number it has recorded in a single year.

In Australia, we are in a literal reading crisis. A 2024 Grattan Institute report said one-third of Australia’s 4 million school children are being failed by an education system that still relies on discredited theories to teach reading. The report warns that students who struggle with reading are more likely to fall behind, disrupt class and later end up unemployed or jailed, with the lifetime cost to the economy estimated at $40 billion.

So when people see AI companies ripping apart books for data, it does not arrive in a vacuum. When reading is already under strain, the sanctity of the written word seems to be pushed further away. I told you! Dystopian as hell.

This article first appeared on PEDESTRIAN.TV

More from Mediaweek

mediaweek
Morning Report

The leading media trade publication in Australia.

Get our top stories straight to your inbox daily by signing up to our Newsletter

By providing your information, you agree to our Terms of Use and our Privacy Policy. We use vendors that may also process your information to help provide our services.