Evgen verzun
Blog
July 30, 2026
Anthropic’s Project Panama and the Data Supply Chain Nobody Is Auditing
This one is slightly off my usual beat, so bear with me. There are no CVEs in this article and nothing here is going to compromise your machine. But it is a story about data provenance, chain of custody, and irreversible operations performed on the only copy of something, and those happen to be things I have spent two decades caring about.
The internet found this story this week and it is furious. Millions of books bought by the pallet, spines sliced off, pages fed through scanners, remains thrown in the bin, all so AI models can read them.
But there are actually two separate stories tangled together here. One of them is documented in court filings and the other is more inference, and the most credible alternative explanation I have read has nothing to do with AI at all.
What Is Actually on the Record
This did not break on Twitter. The news originally came out of Anthropic's copyright case earlier this year, the same litigation that ended in the $1.5 billion settlement.
Filings revealed a dedicated internal effort whose own planning document described the goal as destructively scanning all the books in the world. It ran under a soft codename, and the document stated that the company did not want it publicly known they were working on it. They brought in the former head of partnerships at Google Books to run book acquisitions. They also did buy from used book dealers. One vendor proposal put the target at between half a million and two million volumes inside six months.
The hydraulic cutter is not a rumour either as it is also in the records. Book spines getting stripped so loose pages can be fed through industrial scanners, which is faster and cheaper than imaging a bound book properly, and then whatever is left gets discarded. The contractor they used offers both destructive and non destructive scanning as separate services, so this was not a technical constraint for Anthropic. As far as I have seen, the internal documents never explain why they picked the destructive one. Speed and cost would be the obvious guess.
There is a detail here that explains the appetite. Books printed before roughly 2022 now command a premium specifically because they are guaranteed free of AI generated text. Clean human data is a commodity with a cutoff date after all.
Destruction as a Legal Strategy
The notable fact about the case is that a federal judge ruled the practice was fair use. The reasoning was that Anthropic had bought the copies and then destroyed them, so only one version of each book existed at any given moment. Format shifting, essentially. You owned a thing, then you converted it to another format, but you did not multiply it.
Follow that logic through, though. If you would’ve kept the physical book that would technically mean you copied somebody's intellectual property, which, at scale, starts to look like reproduction. If you destroy it, then you have merely changed the format of your single copy.
Which makes the shredder not just the cheaper option, but also the safer one. There is a straight reading where the destruction is load bearing for the legal argument. Even if everybody finds this part to be barbaric, it is arguably the part that makes the entire operation legally defensible in the first place.
I do not think anyone drafting copyright law imagined it would work this way. It is a decent example of an old framework meeting a situation nobody wrote it for and producing an answer that is technically coherent and intuitively awful.
The Supply Chain Nobody Can See
Here is the point where things start to look familiar to anyone who has audited a vendor chain.
Booksellers across several countries are reporting strange bulk orders and most of them cannot confirm exactly who is buying. According to reporting from 404 Media, there is a book database service that facilitates anonymous bulk orders scaling into the millions attaching an NDA to every engagement and suggesting clients describe the activity as digital preservation. Rare booksellers in the Netherlands have said they are being inundated with bulk purchases they believe come from AI companies, while also saying they cannot actually know.
Anonymity offered as a product feature with non disclosure on every transaction and suggested language to describe what is happening? If you saw that pattern in a software supply chain you would call it laundering and you would not need a second look.
A More Boring Explanation
The most interesting insight I read on this came from a second generation used bookseller in Houston who lived through one of these buying waves back in spring, months before any of this book shredding went viral on Twitter.
His numbers were odd but in line with how others described it. One buyer accounted for roughly 95 percent of his orders on a platform where he normally sells almost nothing. A single order for 70 books, which broke the platform's shipping flow badly enough that he had to call and ask how to fulfil it. All books also shipped to an address unrelated to where the buyer was based. And the most interesting bit is that the titles were completely unremarkable and unrelated, old regional guidebooks and nineties software manuals. It was basically dead inventory.
The bookseller’s conclusion was that it probably wasn’t an AI lab at all. He thought it was arbitrage.
When he checked, every book bought from him was either listed as unavailable on Amazon or priced five, ten, twenty times higher there than what he was charging on the slower platform. So somebody built themselves a little system to spot that gap, buy the cheap side automatically, and route it into Amazon's fulfillment network to sell at a markup. He was careful to say those things and admitted he had no proof of the last part. Just a hunch. But it is a good one that makes for a coherent realistic scenario.
I find that a good deal more plausible than the version where every unusual bulk order on earth is a frontier lab. Which is sort of the point. Once a narrative like this catches on, everything gets attributed to it.
One Copy Is No Copy
Except what our bookseller friend describes lands in roughly the same place where preservation is concerned.
Those books go into warehouse fulfillment. The obscure ones that do not sell within a year or two typically get liquidated as dead stock. If some of them were among the last surviving copies, well, they get gone for good.
Some people go as far as to say that we are living through a second burning of the Library of Alexandria, and the reason nobody notices is that everyone assumes the internet already saved everything. But in reality it didn’t. Somebody has to pay to host that data. When ebook platforms from fifteen years ago folded, they took their exclusive titles with them. Metadata for obscure books is far thinner than people imagine, and when the last physical copy goes there is frequently nothing standing behind it.
Anyone who has done disaster recovery knows the rule. One copy is no copy. You do not delete the source after ingesting it, you do not trust a single medium, and you do not confuse a working copy with an archive. Destructive scanning breaks all three at once, and it does it as a deliberate operational choice on material that in some cases cannot be regenerated.
The Locked Box
The physical book is destroyed after it’s fed into a model. But when you ask the model about it, it will not reproduce more than a handful of words, because copyright constrains the output even though the training was ruled fair.
So the original is gone and the contents are not easily retrievable in any meaningful sense either. The knowledge did not get preserved and it did not get freed for the general public. It got moved into a place nobody can access.
You destroyed the book and locked away its contents, and the thing that survived is a model's ability to write in the general shape of what it read. That is a strange trade to make on the last copy of anything.
The Bottom Line
I am not sentimental about books as objects. I have no attachment to a 1991 word processor manual and I am not going to pretend otherwise. Most of what is being scanned is not precious.
But preserving generational knowledge does matter, and running a business means making hard calls, not defaulting to the cheapest available route every single time. Non destructive scanning exists. It is slower and it costs more, and that is the whole trade off.
We are usually right to be suspicious about how frontier labs source and handle data. This particular story is just not the clean villain arc the thread makes it out to be. Part of it is documented in court records, part of it is people connecting dots that may not connect, and one of the more plausible explanations involves no AI company whatsoever.
The books are not coming back either way, though. That part holds up regardless of which story turns out to be true.