AI Is Destroying Old Books. The Truth Is Strange Enough Without the Conspiracy Theory

FTC Statement: Reviewers are frequently provided by the publisher/production company with a copy of the material being reviewed.The opinions published are solely those of the respective reviewers and may not reflect the opinions of CriticalBlast.com or its management.

As an Amazon Associate, we earn from qualifying purchases. (This is a legal requirement, as apparently some sites advertise for Amazon for free. Yes, that's sarcasm.)

Victorian-style political cartoon depicting books being fed into an industrial guillotine and scanning machine for artificial intelligence.

There are few quicker ways to make people angry than destroying a book. You can throw away an old toaster without anyone accusing you of attacking civilization. Toss a worn-out paperback into the garbage, however, and somebody will probably rescue it before the trash is collected. Books are objects, but we have spent centuries treating them as something more than objects. Burning them is sinister. Banning them is authoritarian. Destroying them feels less like disposing of property than eliminating knowledge.

So when reports began circulating that artificial intelligence companies were buying enormous quantities of physical books, cutting off their bindings, scanning the pages and then disposing of the remains, the Internet reacted about the way you would expect. This time, the horrifying part wasn't made up.

They really are destroying the books.

The best-documented example comes from Anthropic, developer of the Claude artificial intelligence models. Court records arising from copyright litigation revealed an internal effort called Project Panama. Filings later made public show that Anthropic spent tens of millions of dollars acquiring potentially millions of physical books, cutting off their spines, scanning the pages and disposing of the paper copies. The company was building a digital research library for use in developing its models.

One internal planning document put the ambition more bluntly: “Project Panama is our effort to destructively scan all the books in the world.”

That wasn't evidence that Anthropic intended to hunt down the last surviving Gutenberg Bible and feed it through a paper shredder. It was an internal description of a project intended to operate at enormous scale. But the scale really was enormous. The filings describe purchases on a massive scale, with potentially millions of books converted into digital files. The copies used for that destructive scanning did not survive the process. Then booksellers began noticing something else.

Large orders were arriving for hundreds of books at a time. In August 2026, journalists at 404 Media decided to find out where one suspicious shipment was going. Working with a bookseller, they hid a tracking device inside one of the books. It ended up at an Amazon facility in Las Vegas. Further reporting identified the operation as VGT3, an Amazon facility in Las Vegas where books were being destructively scanned for AI training data. An employee later described workers cutting off the spines, scanning the loose pages and dumping the paper into bulk containers afterward.

Somewhere along the way, somebody at VGT3 apparently decided that the appropriate team symbol was a Tyrannosaurus rex holding a book.

You really can't improve upon that.

There is, however, a perfectly rational reason for the apparent barbarism. Books contain something artificial intelligence companies desperately want: enormous quantities of edited, organized, human-written language. Older books also come from before the generative-AI era. If a book was published in 1974, nobody has to wonder whether ChatGPT wrote three chapters of it. The machines want human writing. Libraries are full of it.

And if the goal is to convert millions of bound books into machine-readable files as quickly as possible, cutting off the spine transforms a difficult object into something extraordinarily easy to scan. Instead of photographing two curved pages at a time while protecting a binding, you have a stack of loose, flat sheets that can be fed rapidly through scanning equipment. It's ugly. It's destructive. And it's real.

Advertisement

 

Which is important, because the story doesn't need any help being disturbing. Nevertheless, help arrived. As word of these operations spread through social media, destructive scanning began mutating into something considerably more sinister. AI companies weren't simply destroying copies of old books anymore. They were supposedly buying rare books. In some versions of the story, they were destroying the last surviving copies.

And once those originals were gone, the theory continued, the corporations possessing the digital versions could alter them. History could be rewritten. The evidence of what those books originally said would have disappeared into the shredder.

There are legitimate questions here about preservation, copyright, corporate control of enormous private digital libraries and whether companies consuming books by the millions have any obligation to consider the cultural value of the physical objects they destroy. Those questions are worth asking. But first we need to establish something surprisingly easy to lose sight of when photographs of mutilated books begin circulating online. A book is both a work and a copy of that work.

Destroying one does not necessarily destroy the other. And before we start accusing artificial intelligence of erasing history, perhaps we should find out exactly what is going into the guillotine.

NO, THEY AREN'T BURNING THE LAST COPY

There is a photograph of me somewhere in my mother's collection that probably exists nowhere else. If you destroy it, that photograph is gone. There are also several copies of Stephen King's The Stand in my house. If you destroy one of those, Stephen King has not suddenly written a shorter novel. This distinction should not require much explanation.

And yet it is at the heart of some of the more alarming claims surrounding the books being purchased and destroyed for artificial intelligence.

The language matters. A "rare book" can mean something very specific to a collector. It can also mean an old book that isn't particularly easy to find. And on social media, it can apparently mean any book whose cover looks sufficiently weathered when somebody photographs it next to a story about artificial intelligence. Those are not the same things. The reporting so far establishes that companies have purchased enormous numbers of used books, including older and out-of-print titles. It establishes that those books have been destructively scanned. It establishes that booksellers have received unusual bulk orders large enough to attract attention.

What it does not establish is a systematic campaign to acquire the last surviving copies of books and remove them from existence. In fact, some of the evidence points in the opposite direction. The reporting paints a more complicated picture than the phrase “rare books” suggests. Booksellers caught up in the bulk-buying operation described orders that tended toward ISBN-bearing books rather than antiquarian volumes. A September 2026 New Yorker examination likewise reported that the books being sold into these operations were generally not antiquarian. That does not make every title common, but it is a long way from evidence of a systematic hunt for unique artifacts. That doesn't mean nothing scarce has gone through a scanner.

Advertisement

 

When millions of used books are being purchased, it would be foolish to declare that every one of them was common. An obscure academic monograph might have had a tiny printing. A local history might survive in surprisingly few copies. A forgotten small-press book might technically have an ISBN while only a handful of physical copies remain readily obtainable.

But "some scarce books may have been destroyed" is a very different statement from "AI companies are destroying the last copies of books." And then we get to the second half of the conspiracy. Once the physical book is destroyed, we're told, the company controls the digital version. It can change what the author wrote, alter inconvenient facts and effectively rewrite history because the original evidence no longer exists. Except the original evidence does exist.

It's all the other copies. Books are unusual historical objects precisely because publishing is a technology for making duplicates. With some obvious exceptions—manuscripts, unique documents, annotated copies, limited editions and genuinely rare works—a commercially published book exists specifically because somebody manufactured a lot of substantially identical copies of it. Destroy one copy of a 50,000-copy printing and you haven't destroyed the historical record. You've reduced the surviving population to 49,999.

Nor does scanning a particular copy magically confer authority upon the resulting digital file. If Amazon scans a 1987 history book and changes "Tuesday" to "Wednesday" in its database tomorrow, the copies sitting in libraries, private collections, archives, bookstores and people's basements don't obediently update themselves. Paper has many disadvantages. Remote software updates aren't one of them. This is where the conspiracy theory manages to distract from a much more interesting concern.

The danger isn't that a corporation can destroy one copy of a book and thereby gain control over what the book says.

The danger is scale.

Nobody particularly worries when somebody buys a battered paperback at a library sale for fifty cents and throws it away after reading it. Copies of books disappear constantly through floods, fires, mold, insects, moving-day purges and people who decide that twenty-seven boxes of books are twenty-six boxes too many.

But a company buying and destroying millions of books isn't operating on the scale of somebody cleaning out the garage. At that point, statistics begin to matter. How many copies of each title were printed? How many still survive?

How many are held by libraries? How many exist only in private hands? And, perhaps most importantly, does anybody operating the scanner know the answers before the blade comes down? That is where I become considerably less comfortable.

I don't believe Amazon can rewrite history by cutting the spine off a used book. I do believe that if you're destroying books by the millions, sooner or later you're going to destroy something somebody should have recognized as worth preserving. Not because you're conspiring to erase it.

Because nobody bothered to check.

And that may be a much more believable danger than the conspiracy theory.

WHY DESTROY THE BOOK AT ALL?

Once we establish that nobody has demonstrated a secret campaign to eliminate the last surviving copies of inconvenient books, we're left with a much simpler question. Why destroy them at all? We know how to scan books without cutting them apart. Libraries have been digitizing collections for decades. Put the book in a cradle, photograph the open pages, turn the page and repeat. With enough patience, you can create a digital copy while leaving the original book intact. The problem is contained in three words of that description:

Advertisement

 

Turn the page.

A bound book is a remarkably inconvenient object if your objective is industrial-scale digitization. It doesn't naturally lie flat. Text curves toward the gutter. Pages move. Bindings vary. Some books open easily; others fight you every inch of the way. Somebody—or something—has to turn every page and make sure two pages didn't turn together. Cut off the binding and most of those problems disappear. What was a book becomes a stack of paper. And we have become extremely good at scanning stacks of paper.

Loose pages can move through automatic document scanners at speeds that would be impossible if each page had to be individually turned inside an intact binding. They lie flat. They can be photographed consistently. Software can perform optical character recognition on the resulting images and transform the printed words into searchable text.

The physical object has become inconvenient packaging around the thing the company actually wants.

Data.

If you're digitizing your grandfather's diary, this would be an insane tradeoff. If you're digitizing one book for a historical archive, it would probably be irresponsible.

If you're digitizing millions of commercially produced books because you want the words inside them, the calculation changes dramatically. Time multiplied by millions becomes money. Suppose preserving the binding adds only a minute to the processing of a book. Across five million books, that seemingly insignificant act of preservation has just added more than 83,000 hours of work. And it almost certainly adds considerably more than a minute.

Suddenly the guillotine doesn't look like the tool of a supervillain. It looks like an efficiency measure. That doesn't make it pleasant. It doesn't necessarily make it wise. And it certainly doesn't answer the question of whether every book being processed should be treated as disposable. But it does explain the choice without requiring an evil master plan.

The companies doing this aren't destroying books because destruction is the objective. They're destroying books because preserving them gets in the way of the objective. That distinction matters. It also explains something else about the books being purchased. If your goal is the text rather than the collectible object, condition becomes much less important. A battered ex-library copy with a torn dust jacket contains essentially the same words as the pristine first edition sitting behind glass in somebody's collection.

For the purposes of training a language model, the ugly copy may actually be the smarter purchase.

Buy cheap used copies. Remove the bindings. Scan the pages. Extract the text. Dispose of the remains. Repeat.

Millions of times.

It's less Fahrenheit 451 than Henry Ford.

And that is precisely why the scale deserves scrutiny. An individual collector looks at a book and asks, "What is this?" An industrial process asks, "What's next?" Those are very different questions.

A human being sorting through a box at an estate sale might stop at an unusual volume, look up the title, notice an inscription or realize that the book is considerably less common than the price sticker suggests. An assembly line has every incentive not to stop. Its entire purpose is not to stop. Which leaves an obvious objection hanging over this whole operation.

Advertisement

 

If technology companies can build artificial intelligence capable of digesting the contents of millions of books, surely somebody has invented a machine capable of turning a page. They have. In fact, machines can scan bound books automatically without cutting off their spines at all. So now the question changes.

It isn't whether destroying the book makes scanning easier. Clearly it does. The question is how much easier—and whether that convenience is enough to justify putting millions of books under the blade.

THE MACHINE THAT DOESN'T KILL THE PATIENT

There is an obvious response to everything we've discussed so far. Don't cut the books apart. If artificial intelligence can write computer code, generate photorealistic video and explain quantum mechanics in the voice of a pirate, surely we've reached the technological threshold required to turn a page. We have.

Automated book scanners exist, and some of them look exactly like the sort of machine you would expect to find if a library and a robot had a baby.

One commercially available example is the ScanRobot from Austrian company Treventus. Rather than flattening a book against glass or slicing off its spine, the machine holds the book partially open in a V-shaped cradle. Pages are turned automatically while cameras capture them from above. Software then corrects the distortion created by photographing pages that aren't perfectly flat. Treventus advertises speeds of up to 2,500 pages per hour. At first glance, that would seem to end our discussion.

Problem solved.

Amazon, Anthropic and anybody else with several million books waiting to be digitized can put away the guillotine, buy a warehouse full of page-turning robots and stop murdering the merchandise.

Except there's a reason high-volume document scanners don't come equipped with little robotic fingers.

Turning pages is difficult.

Human beings don't appreciate this because we're extraordinarily good at it. We unconsciously adjust for paper thickness, friction, static electricity, page size and the occasional stubborn sheet that insists on bringing its neighbor along. Machines have to solve all of those problems mechanically. Pages can stick together. Paper can be unusually thin or unusually thick. Old paper can be brittle. Glossy pages behave differently from uncoated ones. Bindings can be tight. Books can contain foldouts, inserts, photographs or other surprises. A machine has to determine that it has turned exactly one page, position that page correctly, capture it and repeat the process hundreds of times without damaging the book it was designed to preserve. And when something goes wrong, somebody has to intervene.

That doesn't make automated nondestructive scanning impractical. Libraries, archives and digitization operations use equipment specifically designed to protect bound materials, and the technology continues to improve. It does make "just use the robot" a little less satisfying than it sounds. The advertised maximum speed of a machine also isn't necessarily the speed of an entire digitization operation. Books still have to be loaded and unloaded. Errors have to be detected. Difficult volumes have to be handled. Images have to be processed and checked. And a machine designed to treat a bound object gently is solving a problem that disappears entirely once you decide you don't care whether the binding survives. That's the brutal advantage of destructive scanning.

Advertisement

 

Cut the spine and you don't need a machine capable of understanding a book. You need a machine capable of understanding paper. We have had very fast versions of those for a long time. For a library preserving a 150-year-old volume, the extra complexity is an easy trade. The continued existence of the physical book is part of the objective.

For a technology company purchasing millions of inexpensive used books solely because it wants the text inside them, preservation becomes another line in the cost calculation. How much does the nondestructive equipment cost? How many operators does it require? How often does it stop?

How many pages can it reliably capture in an hour? How much warehouse space would be needed to store the books after they've been scanned? And perhaps the most corporate question imaginable: Why spend money preserving something we've already extracted everything we believe has value?

That last word is doing a lot of work.

Value.

To the artificial intelligence company, the value may be the words. To a reader, the value may be the book. To a collector, it might be the edition. To a historian, it could be an inscription, a printing variation, a previous owner's notes or some physical characteristic that wasn't important enough to whoever designed the scanning process to record.

A nondestructive scanner preserves all of those possibilities simply by leaving the object alive. Destructive scanning makes a different wager. It assumes we've already decided what matters. And perhaps that's where the argument over these machines has been slightly backward.

The existence of automated page-turning scanners doesn't prove that every $3 used paperback deserves expensive archival treatment. At the scale we're discussing, there may be perfectly reasonable cases where sacrificing an abundant copy is the sensible choice. But the existence of nondestructive scanning does remove one excuse. We cannot say the books have to be destroyed in order to digitize them.

They don't.

We're choosing to destroy them because, under certain circumstances, it's faster, easier or cheaper than preserving them. Which means the real question isn't whether we possess the technology to save the book. It's whether we think this particular book is worth saving. And somebody has to make that decision before the blade comes down.

WHEN DOES AN OLD BOOK BECOME A RARE BOOK?

Here's where I start sympathizing with the people yelling about rare books. Not because I think artificial intelligence companies are systematically hunting down the last surviving copies so they can rewrite history. Because I don't trust anybody to know which copy is the last one. Publishing produces numbers that can create an illusion of permanence.

Print 100,000 copies of a book and 100,000 sounds like forever. Surely some of those copies will always exist. Except books aren't immortal. They get thrown away. They get wet. They get moldy. They get eaten by insects. Libraries discard them. Bookstores pulp unsold inventory. People die and their children look at six bookcases containing Dad's collection and decide they need the room. Do that for fifty years and nobody really knows how many copies remain.

Advertisement

 

Libraries can tell us what they hold. Dealers can tell us what they have for sale. Collectors know what's sitting on their own shelves. None of those numbers tells us how many copies still exist. That matters because "rare" is an extraordinarily slippery word. A first edition of a famous nineteenth-century novel may be rare in the traditional collectible sense while the text itself exists in millions of later editions, reprints and digital copies. Destroying that first edition would be an act of cultural vandalism, but it wouldn't erase the novel.

An obscure technical manual printed in 1973 might be worth twelve dollars. It could also be considerably harder to replace. Price and scarcity aren't the same thing. Neither are age and importance.

A forgotten regional history, university-press monograph, church anniversary book, small-press memoir, industrial handbook or proceedings from some long-defunct professional organization may never become valuable enough for collectors to fight over it. That doesn't mean nobody will ever need it. Historians are professional scavengers. The mundane documents everybody ignored at the time are frequently the things somebody desperately wants decades later. What did people actually believe?

How did this company describe its manufacturing process? What businesses existed on this street? What terminology did doctors use? What did this organization tell its members?

Sometimes the answer is sitting inside a book nobody thought was important enough to save. This is where industrial-scale destructive scanning creates a legitimate preservation problem even if every conspiracy theory surrounding it turns out to be nonsense. The individual decision can be perfectly rational. Here's a used book. It has an ISBN. We bought it legally. Another copy is listed online for $7.95. Cut it.

Then do that a million times. At some point, probability starts working against you. Sooner or later the conveyor belt is going to receive something scarcer than anybody realized. Maybe not the last copy on Earth. Maybe not even the last ten. Just something for which the surviving population was already moving in the wrong direction.

And there's an additional irony. The very books most attractive for artificial intelligence training may include material that isn't readily available online. That's part of their value. If everything in a book already existed as clean, accessible digital text, there would be considerably less reason to buy a physical copy, cut it apart and scan it.

So we're potentially destroying physical copies partly because their contents haven't already been adequately preserved digitally. That should at least make somebody hesitate. Fortunately, this doesn't require an international Ministry of Book Preservation deciding whether your 1994 Windows manual qualifies as a cultural treasure. We already have tools that could help.

Check library holdings. Check the edition and printing. Check availability in the used-book market. Flag unusually scarce titles. Identify older books and small print runs for additional review. Route questionable material to nondestructive scanning.

And when in doubt, don't cut it.

For companies building some of the most sophisticated information-processing systems humanity has ever created, determining whether another copy of a book exists seems like a problem that ought to be within reach. There will still be judgment calls. I don't expect anyone to preserve every battered copy of The Da Vinci Code because future historians might someday need one. There are enough copies of certain books in circulation that sacrificing one for digitization isn't meaningfully different from losing one to a flooded basement.

Advertisement

 

But "probably another one somewhere" becomes less comforting as the number of books being destroyed climbs into the millions. That's the part of this story worth worrying about. Not a secret campaign to make inconvenient books disappear. Not a digital Ministry of Truth waiting to change yesterday's words.

Something far more ordinary. We built a process designed to move very quickly through an enormous pile of books. And moving quickly is exactly how you fail to notice the one you should have stopped for.

SAVE THE BOOK. LOSE THE CONSPIRACY.

So where does that leave us? Artificial intelligence companies really are buying physical books by the millions. They really are cutting the bindings off at least some of them. They really are scanning the pages.

And the copies subjected to that process really are being destroyed. None of that requires scare quotes. None of it needs to be exaggerated to make it interesting. What the evidence does not establish is the story that has grown around those facts: that technology companies are deliberately seeking the last surviving copies of rare books, destroying them so no physical evidence remains, and positioning themselves to alter the digital versions and rewrite history. That's a much better movie.

It's also a distraction from the questions we should actually be asking. Should companies assembling enormous private digital libraries be permitted to destroy millions of physical books without making any effort to determine their scarcity?

Should books that appear to have limited surviving copies automatically be diverted to nondestructive scanning?

Should the resulting scans eventually become available to libraries or researchers, particularly when the physical copy used to create them no longer exists?

And when a corporation decides that the only valuable thing inside a book is its text, who gets to tell it that the book itself might also contain information worth preserving? Those aren't conspiracy theories. They're policy questions. They're preservation questions.

And they're questions we're going to have to answer more frequently as artificial intelligence develops an appetite for information that still exists primarily in physical form.

There is also something slightly uncomfortable about watching people suddenly become horrified by the destruction of books.

We have been destroying books forever.

Publishers pulp unsold copies. Libraries weed collections because shelves are finite. Bookstores dispose of inventory they cannot sell. Schools replace old textbooks. Families empty houses after relatives die. Water, fire, insects and neglect take care of countless others.

Nobody launches a national campaign every time a box of Reader's Digest Condensed Books goes into a dumpster. What has changed is who is destroying them and why. When Grandma throws away a box of old paperbacks, we understand the transaction. She doesn't want them anymore.

When an artificial intelligence company buys the same books, cuts them apart, extracts the words and throws away the paper, something about it feels different. Because the company did want the books.

It wanted them very badly.

It simply didn't want them to remain books.

I understand why that bothers people. It bothers me a little, too. Given the choice between destroying a physical book and preserving it, I'd rather preserve it. If automated scanning technology can keep a book intact at a reasonable cost and speed, I'd prefer companies use it. And when we're talking about obscure, old or potentially scarce material, I'd go further: somebody should be checking before anything irreversible happens.

Advertisement

 

But preservation arguments don't become stronger when we attach claims to them that the evidence doesn't support.

Quite the opposite.

Tell people that millions of books are being purchased and destroyed to feed artificial intelligence systems and you've already given them something worth thinking about.

Tell them the last copies are being secretly eliminated so corporations can rewrite the contents later, and you've handed everyone who wants to dismiss the legitimate concern an easy escape hatch. They no longer have to discuss preservation. They only have to laugh at the conspiracy theory. That's a shame, because there is something genuinely strange happening here.

For most of the modern era, digitizing a book was understood as an act of preservation. We made another copy so the information could survive if something happened to the original. Now we've reached the point where, at sufficient scale, digitization can mean the opposite. We destroy the original because we've decided the copy is enough. Maybe sometimes it is.

Maybe sometimes the paperback really is just one of 100,000 interchangeable containers for the same words, and sending one through a scanner is no greater cultural tragedy than sending it through a recycling bin. But sometimes it won't be. And when you're processing millions of books, "sometimes" eventually becomes somebody's irreplaceable mistake.

That's enough reason to demand better judgment, better safeguards and, where appropriate, better technology. We don't need to imagine artificial intelligence erasing the past to justify preserving it. If you have to invent a conspiracy to make people care that we're feeding millions of books into guillotines, you've somehow managed to make an already fascinating story less interesting.