Strip, scan and shred: AI companies destroying rare books for profit

In July 2026, Senator Josh Hawley (R-Missouri) chaired a Judiciary subcommittee hearing focusing on Big Tech’s role behind the piracy of copyrighted content to fuel companies’ artificial intelligence (AI) models. The hearing featured witness testimony from bestselling author David Baldacci, AI experts and law professors.

Hawley’s claim – AI companies had crossed the line of technological innovation into corporate crime.

“This hearing is about the largest intellectual property theft in American history,” Hawley told the hearing. “AI companies are training their models on stolen material, period. And we’re not talking about these companies simply scouring the internet for what’s publicly available. We’re talking about piracy,” Senator Hawley said.

“Are we going to protect our creative community, or are we going to allow a few mega-corporations to vacuum it all up, digest it, and make billions of dollars in profits – maybe trillions – and pay nobody for it,” Hawley said. “That’s not America.”

AI corporations ain’t no dummies.

Litigators and lawmakers are accusing us of theft? Hell, we’ll teach him. Instead of stealing the copyright, we’ll buy the physical books, scan them and shred them. That’ll teach ‘em.

With the Office of Technology Assessment in mothballs, Congress has yet to lift a finger to confront what one British broadsheet labeled “cultural barbarism” – the physical destruction of the world’s books.

For centuries, books have represented one of humanity’s most durable technologies. They require no electricity. They do not need a subscription. They can survive war, political upheaval, and the collapse of empires. A book printed hundreds of years ago can still carry the thoughts of someone long dead into the hands of someone not yet born.

But in the age of artificial intelligence, some of the world’s largest technology companies see books differently. They see data.

The rise of large language models has created a new appetite for written material on a scale never seen before. AI companies need enormous quantities of high-quality text to train systems capable of generating essays, legal arguments, computer code, and human-like conversation.

And increasingly, they are looking beyond the internet. They are looking at bookshelves. According to 404 Media and other outlets, AI companies and their vendors have begun purchasing large quantities of physical books, scanning them, and in some cases destroying the original copies afterward.

The practice has raised questions among authors, historians, booksellers, and lawmakers about what happens when humanity’s cultural archive becomes training fuel for private artificial intelligence systems.

The controversy intensified after court filings in a case involving Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson against Anthropic PBC, a copyright lawsuit in federal court in California, revealed details about the company’s efforts to acquire and digitize books for AI training. (Last year, Anthropic settled the case for $1.5 billion)

The project, internally referred to as Project Panama, involved efforts to obtain large quantities of books and convert them into searchable digital data.

One phrase from internal corporate documents attracted particular attention: “We don’t want it to be known that we are working on this.”

The reason for secrecy was not necessarily the scanning itself. Digitizing books has existed for decades. Libraries, universities, and companies have all undertaken similar projects.

The controversy was the scale. And the destruction.

AI companies argue that books provide some of the highest-quality training materials available. Unlike much of the open internet, which contains spam, misinformation, and increasingly AI-generated content, books are edited, structured, and often written by experts in the field.

A company called ISBNdb, which provides book data services, described pre-large language model books as particularly valuable because they were created before the internet became filled with AI-generated text. “Physical books published before this date are structurally clean of modern poisoning tools,” the company reportedly stated.

For AI developers, the appeal is obvious. Books contain history, science, literature, legal analysis, philosophy, technical manuals, and countless forms of human reasoning. They provide models with examples of how humans organize ideas. But acquiring millions of books presented a problem. Publishers own copyrights. Libraries have restrictions. Authors have objected to their work being used without permission.

According to court documents, Anthropic initially relied on digital copies obtained from online sources – including pirated repositories. The company later faced copyright litigation over its acquisition and use of books.

The legal fight pushed the company toward another strategy: acquiring physical copies. Rather than asking every publisher and author for permission, companies could purchase books through legal marketplaces and distributors.

Under the first-sale doctrine, a person who legally purchases a physical copy of a copyrighted work generally has broad rights over that individual copy.

The argument made by AI companies was that while downloading unauthorized digital copies might be illegal, scanning purchased books into digital formats for internal AI training was “transformative fair use” and therefore protected under fair use. The destruction of the physical copy became a central part of the argument.

The process described in court documents was industrial. Books were purchased in large quantities. Their spines were removed. Pages were separated and processed through high-speed scanners. The resulting digital files became searchable collections of text. The physical books were then discarded, shredded or recycled.

Supporters of the practice argue that this is comparable to other forms of digitization. Libraries have long scanned materials to preserve and analyze them.

But critics argue that the scale and purpose are fundamentally different. A library scans a book to preserve knowledge and make it available. An AI company scans a book to build a commercial product. That distinction matters. The original book may disappear, while the resulting digital copy remains locked inside a private corporate system.

The public loses access to the physical object. The company gains a competitive advantage – forcing individuals to go to these AI companies instead of public resources to gain access to certain materials.

The effects are being felt by booksellers. One used bookseller interviewed by 404 Media described a sudden increase in purchases. Instead of selling roughly 20 books per week, sales jumped into the hundreds.

The orders were unusual. Customers were buying random collections of books, often identified through ISBN numbers.

The seller suspected AI companies were behind the purchases. “I personally have mixed feelings about all of this,” the bookseller said to 404 Media.

The sales helped financially and cleared inventory that otherwise might have been left unsold. But the seller expressed concern about rare and out-of-print books being destroyed.

“I don’t like the end-use, and I don’t like that uncommon books are being pulped,” the seller said.

That concern has spread among rare book dealers internationally. Some European booksellers have reported receiving large purchase requests for thousands of titles – often across multiple languages.

The fear is not necessarily that AI companies will buy every valuable book. It is that they may unknowingly destroy irreplaceable copies.

A mass market novel may have millions of copies. A technical manual from decades ago may have only a handful. A local history book may exist in a few archives. Once those copies are destroyed, they cannot easily be recovered.

The most unusual aspect of the debate is that destruction itself may strengthen the legal argument. If a company buys a book, scans it, and then continues selling or distributing copies, copyright concerns are obvious. But if the company destroys the original after creating a digital copy, it can argue that it is not creating a competing marketplace for the book.

The original copy is gone. The digital version is used internally.

In Bartz v. Anthropic, U.S. District Judge William Alsup concluded that scanning lawfully purchased books for AI training could qualify as fair use, while allowing plaintiff’s claims based on Anthropic’s earlier acquisition of pirated books to proceed.

Critics argue this creates a strange incentive structure. The more destructive the process becomes, the stronger the company’s fair-use argument may appear.

A book preserved in an archive remains a public cultural object. A book converted into private training data becomes a corporate asset.

The controversy over books is part of a broader debate about artificial intelligence.

Supporters of rapid AI development argue that access to enormous amounts of information is necessary for technological progress. They compare AI training to human learning: people read books, absorb information, and create new ideas.

Critics argue that AI companies are extracting centuries of human knowledge while concentrating the benefits among a small number of corporations.

The debate is not only about copyright. It is about ownership. Who controls the accumulated knowledge of civilization? Should humanity’s written record become the private foundation of commercial AI systems? Or should there be public rules governing how that information is collected and used?

These questions are increasingly moving from technology circles into government. Lawmakers are beginning to examine AI’s impact on copyright, labor, intellectual property, and public information.

The fight over books may become one of the first major battles over what happens when artificial intelligence consumes the products of human culture.

For thousands of years, books survived because societies believed they were worth preserving. The AI era is forcing a new question: What happens when the most valuable part of a book is no longer the book itself – but the information inside it?

And who gets to decide what is worth preserving?