← All Guides

Book News

Project Panama: Why an AI Company Destroyed Millions of Books


On the Apple TV+ show Silo, there's a word for the objects the government doesn't want you to have: relics. Old books, old photographs, anything left over from the world before. Keeping one is against the law — not because the ideas inside it are dangerous exactly, but because the people in charge have decided that ordinary citizens don't get to decide that for themselves. An entire department exists to make sure relics stay confiscated, cataloged, and out of circulation.

It's a good piece of dystopian fiction. It is not, as far as anyone has reported, what happened at Anthropic.

But this year, a stack of unsealed court documents described something that rhymes with it in an uncomfortable way — and the differences turn out to matter more than the similarities. Here's the story, as carefully as we can tell it, with the parts that are confirmed kept separate from the parts that aren't.

What Project Panama actually was

Anthropic is the company behind the Claude chatbot. Starting in early 2024, it ran an internal program with the code name Project Panama. An internal planning memo — later unsealed as part of a copyright lawsuit — described the goal in one blunt sentence: “Project Panama is our effort to destructively scan all the books in the world.”

The same memo explained why it was a “soft codename” at all: “We use ‘soft codename’ because we don't want to be known that we are working on this.”

Here's how the program worked, according to the reporting that followed the document unsealing. Anthropic bought used books in bulk — tens of thousands of copies at a time — mostly through two ordinary secondhand booksellers, Better World Books and the UK's World of Books. A separate vendor then handled the destructive part: a hydraulic cutting machine sliced the spine off every book, so the loose pages could be fed through high-speed industrial scanners. Once the text was captured, the paper went into recycling.

One vendor listing described the target scale plainly: converting 500,000 to 2 million books over a six-month period. Reporting across multiple outlets consistently describes the purchases, over the life of the project, as being in the millions.

This wasn't hidden in a warehouse somewhere overseas doing something exotic. It was ordinary used bookstores, an industrial scanning vendor, and a lot of paper.

Why older books, specifically

If you're wondering why an AI company would go to this much trouble instead of just training on text already available online, the internal reasoning that's been reported is fairly candid: Anthropic wanted text that could teach Claude, in the memo's words, “how to write well,” rather than mimicking “low quality internet speak.”

There's a broader industry logic behind that, too, and it's worth spelling out because it's not unique to Anthropic. Books published before the recent explosion of AI-generated web content are, by definition, clean human-written text. As more of the open internet fills up with AI-written filler, AI companies increasingly worry about something researchers call model collapse — the slow degradation that happens when a model gets trained on text that was itself written by a model, compounding small errors and flattening style generation after generation. Older, human-authored books are a hedge against that. “Pre-2022” has become industry shorthand for text that predates the AI content flood, even though it's a rough boundary, not a precise cutoff any company has put a single confirmed date on.

The honest correction: what's confirmed, and what isn't

This is the part of the story that's easiest to get wrong, so we want to be direct about it.

Some coverage of Project Panama claims that rare or first-edition books were specifically targeted and destroyed. We're not going to repeat that claim, because it isn't confirmed. Snopes examined the underlying documents and rated the broader claim about industrial-scale destructive scanning “Mostly True” — but explicitly flagged the rarity question as unresolved, noting the internal documents refer to targeting “less common” titles without making clear what that means for a book's actual value or rarity in the collectibles market. Anthropic, for its part, has denied specifically targeting rare or valuable collectible books.

So here's the line we're drawing, and staying on: confirmed — destructive scanning, at industrial scale, of ordinary used books, sourced mainly through mainstream secondhand booksellers. Not confirmed — that any specific rare, first-edition, or collectible book was destroyed. If you see a version of this story that treats the rare-books claim as settled fact, treat that version skeptically. We'd rather tell you what we don't know than borrow a more dramatic story that isn't fully backed up.

We're also not going to name specific titles that supposedly went through the cutter, because no reporting we can find actually identifies any. If a version of this story does, ask where that title came from.

The lawsuit, and the ruling that actually matters

Project Panama came to light because authors sued. Andrea Bartz (We Were Never Here), Charles Graeber (The Good Nurse), and Kirk Wallace Johnson (The Feather Thief) brought the case — Bartz v. Anthropic — that eventually forced these internal documents into public view.

In June 2025, U.S. District Judge William Alsup issued a ruling that's now treated as a landmark decision on how copyright law applies to AI training. It was a split decision, and the split is the part worth understanding, because it's where people most often get the story wrong.

Judge Alsup ruled that training an AI model on lawfully acquired book text is “exceedingly transformative” and therefore fair use. He also ruled that destructively scanning a book you legally purchased — converting a physical copy you own into a digital one, even destructively — is also fair use, treating it similarly to format-shifting something you already own.

But that ruling did not extend to a separate part of Anthropic's operation: millions of pirated ebooks pulled from shadow libraries, including sources like LibGen and a mirror known as the Pirate Library Mirror. That piracy was not covered by the fair-use finding. It's specifically what led Anthropic to settle.

The settlement, approved by the court in mid-2026, came to $1.5 billion — the largest publicly reported copyright settlement in an AI case to date. The money went to the class of authors and rightsholders whose books had been pirated, averaging roughly $3,000 per book — a small fraction of the $150,000-per-work maximum statutory damages available under copyright law, but a record-setting total given the scale of the pirated library involved.

Put plainly: buying a book, destroying it, and training on the text was ruled legal. Pirating the same text instead was not, and that distinction is worth about a billion and a half dollars.

Anthropic's own public position, via a company deputy general counsel, was that the settlement addressed “how some materials were acquired, not whether we could use them” — the company isn't conceding that AI training on books is unlawful, only that piracy was the wrong way to acquire some of it.

What it means for readers and the used-book market

If you're a reader, the honest answer is that the books already on your shelf were never part of this specific program. Nobody is coming for your paperbacks. But the story does point at something real happening underneath the used-book trade: AI companies have become a genuine source of institutional demand, buying at a scale that didn't exist in this market five years ago.

Booksellers are starting to notice. One Dutch bookseller described receiving a request for roughly 3,000 copies of a single title and initially assumed it was spam or a phishing attempt — before realizing it was a legitimate, if strange, bulk order tied to AI training. That's not an isolated anecdote; it's an early signal of a new kind of buyer showing up in a market that used to be mostly readers, libraries, and collectors.

Whether that's good or bad for the used-book trade long-term is genuinely unclear. More institutional demand could mean better prices for sellers moving large backlist inventory. It could also mean certain categories of books — the “less common” titles the internal documents referenced — get harder to find secondhand as bulk buyers compete with individual readers for the same stock. It's too early to know which effect wins.

Watch the video version

We also made a video walking through the whole story — the memo, the cutter, the settlement, and what's rumor versus record:

One honest thing about us

We want to say this plainly rather than bury it in fine print: this article, and the video it's paired with, were produced with help from AI — the same broad category of technology this story is about. We're not reporting on Project Panama from some position outside the AI industry looking in. We're a small media operation that uses AI tools in production, writing about a much larger AI company's decisions regarding the raw material — books — that trains those tools in the first place. We think that proximity is worth naming instead of pretending it isn't there.

Why the object still matters anyway

Here's where the Silo comparison we opened with actually breaks, in a way that we think is the most interesting part of this whole story.

In Silo, relics are outlawed to keep ideas contained — to control what people are allowed to know. That is not what happened here. Nobody at Anthropic tried to erase the ideas in these books. If anything, the opposite: they wanted the words badly enough to digitize every single page before doing anything else. What they treated as disposable wasn't the text. It was the object — the specific paperback, the specific copy, the one with a particular water stain or a particular previous owner's name written inside the cover.

That's a stranger kind of loss than censorship, and in some ways a more modern one. It's not a story about ideas being suppressed. It's a story about a physical object being judged worth less than the information that can be extracted from it — data that model can now hold, cheaply, forever, distributed across a company's servers.

Which is exactly why the object still matters, on purpose, even in a world where the words themselves can live somewhere else entirely. A language model can hold every sentence a book ever contained. It cannot hold the specific copy that got dog-eared on a road trip, or handed down from a parent, or picked up for a dollar at a library sale and read on a rainy weekend. That copy means something because it happened to be yours, not because of the information it carries. That's not data. It's closer to a relic — the good kind, the kind worth keeping on purpose.

If Project Panama proves anything beyond the legal questions it raised, it might be that the difference between a book and its text is not a technicality. It's the whole point of owning one.


If this made you want to hang onto the books already on your shelf, we made a set of free printable bookmarks — nothing to buy, just print, cut, and use. Find them at hookedtobooks.tv/bookmarks.

Related Reads