The Book Is Gone;
Where's the Scan?
Here we are again with the algorithmic outrage machine, and I’m noticing the manufactured nature of it all. Last month it was “data centers use too much water” with zero nuance, and it caught virality. Too much water compared to what exactly? The major offenders are agriculture (which we need) and lawn maintenance (which is debatable). I’ll agree on this: we probably shouldn’t be installing data centers where they can’t cool efficiently, like Texas. The tax incentives push businesses to relocate there, and required resources seem to be an afterthought. And this won’t be the only invisible hand I mention in this post.
This Cycle’s Headline
Today it’s “AI Company Destroys Two Million Books”, with an X post as ground zero. Sounds pretty bad, and the headline is all most people read. Indeed, the term “18th-century botanical text” seems to be trending, “Books that outlasted wars, fires, and centuries of handling…”, as X users are assuming these are the targets of the AI trainers, text that would be in pre-1931 public domain (no need to destroy it). However when we look at the nuance, it’s almost a cartoon villain level of framing that tries to hide what the books actually are. Charlie Becker shed the light we needed here:
“Every ingredient of the viral story is true. ISBNdb really does advertise sourcing books for AI companies. Anthropic really did destructively scan millions of books though they were mostly acquired through things like library deaccessions, not bookstore inventory. And booksellers really are getting bizarre bulk orders. But “rare” here isn’t what you’re picturing. It’s The Insider’s Guide to Metro Denver (1995) and how-to-use-WordPerfect 1991 manuals. Obscure, but often not precious.”
Rare Does Not Mean Valuable
So Anthropic acquired from library deaccessions, not bookstores. And ‘rare’ doesn’t necessarily mean ‘valuable’. Any single viable item was likely pulled individually by the seller, not sold in a bulk lot. Becker’s conclusion, that these new waves of orders are not AI companies themselves, but from AI-driven Amazon reselling arbitrage, makes a lot more sense as an inciting incident. The deep reporting on Bartz v. Anthropic landed six months ago, on a ruling where the judge held the destruction of the original material was actually what protected Anthropic’s position. That’s why it feels so manufactured to me: why now, six months later, is this a viral story? That ruling is the invisible hand; another AI company looking at this case will make book destruction the standard operating procedure to keep out of legal trouble. It’s news now because of the incredible capabilities of models releasing and people using it for ‘this’ to make a buck.
Keep the Scan
I think Becker’s work is admirable, and I support it. But I don’t think it’s the entire solution. What should happen is in the interest of what everyone is actually mad about: irreversible, un-auditable destruction. The last surviving copy of a book shouldn’t be tossed without preservation, even if its content is not valued at the present. Anthropic ‘should’ be preserving these texts, uploading quality scans to the Internet Archive. Instead they’ll likely remain inside the training vaults, as the judge encouraged. I’m with Becker on this one; this isn’t necessarily the AI company boogeyman, and the scapegoating is protecting the real offenders.
Films and Games Are Next
It’s a preview of what’s to come. Text is the corpus now and it’s increasingly becoming a well run dry. I’d bet the next archival pressure to come down will be on films and video games as this continues. Games are dying ‘right now’ as MMOs get pulled, projects canceled, IP licenses not renewed. After the company shuts down, that software history and collective human achievement are potentially lost forever. Preservation is going to be everything as this data becomes more valuable, and again we have the hand of the law pushing in the direction ‘away’ from proper archival, influenced by lobbyists.
The Copies Paradox
It’s funny to me; you can make a thing, copy it near-infinitely, and sell the copies. But don’t you dare exercise any control of a copy you ‘purchased’, unless you have a proper license that cannot realistically be obtained. That kind of stance makes real archiving nearly impossible. You can see this with the death of physical media as the producers try to muscle out resellers, leaving rights-holders with less value than before.
The laws are not keeping up with the times, and when they do, it favors big players with fat wallets. You are not the beneficiary.
Time to Change Course
The dedicated archivists have a tough road ahead, but this moment now is the time to change course. The pressure on AI companies to archive properly is good, it only lacks the nuance it needs to resonate with everyone. The Amazon arbitrage players may fail, leaving us with no preservation whatsoever on those texts. What’s lacking now is the pressure on the law to reflect the will of the people, and not corporations who see archiving as a threat to shareholders.
Sources
The July 2026 viral wave:
- @HedgieMarkets X post (Jul 26, 2026; the post that resurfaced the story)
- Futurism: “AI Companies Are Reportedly Destroying Rare Books” (Jul 2026)
- AP via WBAP: “The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions” (Jul 27, 2026)
Charlie Becker (bookseller’s view):
- X post: “Every ingredient of the viral story is true…”
- Substack: “Is an AI Company Buying Up All the World’s Books?” (the arbitrage analysis, with sales data)
Bartz v. Anthropic:
- Washington Post: “Anthropic ‘destructively’ scanned millions of books” (Jan 27, 2026; the unsealed-docs deep dive)
- Business Standard: how Anthropic ended up paying $1.5B for training on pirated books (Sep 2025 settlement; the pirated corpus, distinct from the purchased-and-scanned one)
Preservation and archives:
- Harvard Institutional Data Initiative (roughly 1M public-domain books released as an open training dataset)
- Nieman Lab: news publishers limit Internet Archive access over AI scraping concerns (Jan 2026)
- EFF: “Blocking the Internet Archive Won’t Stop AI” (Mar 2026)
Games and physical media:
Frequently Asked Questions
Is an AI company buying up all the world's used books?
Probably not, at least not the wave that went viral in July 2026. Bookseller Charlie Becker, whose family runs Houston's largest used bookstore, traced the bulk orders in his own sales data: roughly 95% of orders came from a single buyer, and the targeted titles had Amazon listings that were out of stock or priced 5 to 20 times higher than his copies. That is the fingerprint of Amazon FBA resale arbitrage, not AI training acquisition. Anthropic's confirmed destructive scanning sourced books mostly through library deaccessions, not bookstore inventory.
Did Anthropic really destroy millions of books?
Yes. Court records from Bartz v. Anthropic and follow-up reporting describe the pipeline: buy used books in bulk, cut the bindings, scan them into internal machine-readable files, and discard the paper. In June 2025 Judge Alsup ruled that buy-scan-destroy pipeline was fair use. The $1.5 billion settlement in September 2025 covered a separate corpus of pirated ebooks, not the purchased-and-scanned books.
Are the books AI companies destroy rare or valuable?
Mostly neither. The inventory in the viral reports is liquidation-tier stock: 1995 regional tourist guides, early-90s software manuals, foreign-language overstock. Rare does not mean valuable. The legitimate concern is that some obscure titles may be last surviving copies, and because the bulk orders are anonymous with no manifest, nobody can audit whether a last copy was destroyed.
Do AI companies preserve the scans of the books they destroy?
There is no public record of any AI lab depositing destructive-scan output with the Internet Archive or any library. The legal incentive points the other way: the fair-use ruling credited Anthropic for keeping the scans locked in a private corpus rather than redistributing them, so the physical copy is destroyed and the digital copy stays in a training vault.
What should happen to the scans instead?
Deposit them with an archive. The mechanism already has a proof of concept: Harvard's Institutional Data Initiative, funded by AI companies, released roughly one million public-domain books as an open training dataset. The same shape could extend to in-copyright scans as a dark archive, with sealed deposits released when copyright expires or rights are cleared, so destroying the paper no longer means losing the text.
What happens to bulk-bought books that resellers cannot sell?
They get liquidated or pulped. That is the quieter black hole: a failed arbitrage play destroys the book with no scan and no surrogate, which is worse for preservation than the AI pipeline, which at least leaves a digital copy (locked up, but existing). No AI company is required for that loss.