This Week in AI
  Anthropic's Project Panama: Books Destroyed to Train Claude  Microsoft AI For Beginners Course Tops GitHub Trending  Claude Opus 5: Anthropic's Model for Long-Running Agents  GPT 5.6: OpenAI Pushes the Price-Performance Frontier  Hyperagent: No-Code AI Agent Platform From Airtable Founder  Chinese AI Models Overtake US Rivals in Global Usage  6 Trending Open-Source AI Repos on GitHub This Week  AI 2040 Plan A: The Case for a Frontier Pause
LLM Launches & Updates

Anthropic's Project Panama: Books Destroyed to Train Claude

Anthropic's secret Project Panama bought and destructively scanned physical books to train Claude — after a $1.5B payout over 7 million pirated books.

Anthropic's Project Panama: Books Destroyed to Train Claude

> **TL;DR:** Anthropic ran a secret program, reportedly dubbed "Project Panama," that bought massive quantities of physical books — antiques included — cut off their spines for high-speed scanning, recycled the remains, and used the text to train Claude. The pivot followed a $1.5 billion payout over more than 7 million pirated books the company had used earlier. Elon Musk's team reportedly takes the opposite approach, scanning rare books non-destructively and preserving them in a library.

Key Takeaways

- Anthropic's secret "Project Panama" bought huge quantities of physical books — including antiques — cut off their spines for rapid scanning, and recycled the destroyed copies, per a Washington Post report. - Owning print copies sidesteps ebook licensing limits; the scanned text trained Claude and the scans were never redistributed. - The program followed a $1.5 billion payout over more than 7 million pirated books Anthropic had downloaded for pre-training. - Elon Musk's team reportedly preserves rare books in a library and scans them non-destructively while still using the text for AI training. - Books are the AI industry's scarcest high-quality data source — expect more of the printed record to be industrialized into training tokens.

Anthropic bought physical books in enormous quantities, sliced off their spines, ran the loose pages through high-speed scanners, and recycled what was left — all to produce training data for its Claude models. The secret program, dubbed "Project Panama" and revealed in a Washington Post report, consumed antique volumes along with everything else. It is the clearest picture yet of how far a frontier lab will go to secure high-quality text, and it follows directly from one of the most expensive data mistakes in the industry's short history: a $1.5 billion payout over the more than 7 million pirated books Anthropic used before it started buying its own.

What Is Project Panama?

Project Panama was a secret Anthropic effort built around a blunt physical premise: purchase massive numbers of printed books, destroy them for rapid digitization, and feed the resulting text into Claude's training corpus. That is the picture that emerges from the Washington Post's reporting, which describes an operation treating the printed word as bulk industrial input — including antiques, not just recent paperbacks.

The pipeline: cut, scan, recycle

The mechanics explain the destruction. Cutting the spine off a book converts a bound volume into a loose stack of pages that can be fed through an industrial sheet scanner at high speed — dramatically faster than imaging a fragile, bound book page by page. Once a book was digitized, the paper remains were recycled. The scans themselves stayed inside the company: Anthropic used the text to train Claude without redistributing the copies it made, an important distinction in how the program was structured.

![A code editor interface is shown with a terminal window open, displaying a command line prompt and a snippet of code](https://supabase.srv1729373.hstgr.cloud/storage/v1/object/public/blog-images/speka-info/anthropic-project-panama-book-scanning-1-664732892ded5545.png)

Why Buy Books Only to Destroy Them?

The short answer is ownership. Ebooks are generally sold under license agreements that restrict copying and bulk text extraction, which makes them awkward raw material for model training. A physical book is different: buy it, and that copy is yours. By purchasing print editions outright, Anthropic sidestepped the licensing limits attached to digital books entirely.

Destruction, meanwhile, is about throughput rather than ideology. Careful, preservation-grade scanning of a bound book is slow and expensive. Guillotine the spine and the same book becomes scanner-ready in seconds. When the goal is digitizing books at the scale a frontier LLM demands, the destructive route wins on speed and cost — and the recycled remains are simply what is left over when literature is processed as feedstock.

The antiques are the detail likely to dominate the debate. A mass-market paperback pulped after scanning is a curiosity; an antique book cut apart and recycled is, for many readers, a small cultural loss that no amount of digitized text repays.

The $1.5 Billion Mistake Behind the Pivot

Project Panama did not appear out of nowhere. Before it began buying books, Anthropic downloaded more than 7 million pirated books to pre-train its models. That shortcut ended in a major copyright lawsuit and a $1.5 billion payout — and that outcome is what pushed the company toward purchasing physical copies instead.

Seen in that light, the economics of the program are straightforward. Whatever it costs to buy, cut, and scan warehouses full of used and antique books, it is hard to imagine the bill approaching what the piracy settlement cost. The lawsuit effectively reframed legitimate acquisition — even at absurd physical scale — as the cheap option. That may be the most consequential part of this story for the wider industry: the price of taking copyrighted text without permission is now concrete, and it is large enough to change how the biggest labs behave.

![An antique book with a worn leather cover and frayed pages, held by a person's hand](https://supabase.srv1729373.hstgr.cloud/storage/v1/object/public/blog-images/speka-info/anthropic-project-panama-book-scanning-2-7d30c14f9b356cae.png)

Elon Musk's Team Took the Opposite Approach

Destructive scanning is a choice, not a necessity — and the reported counterexample comes from Elon Musk. Musk has reportedly instructed his team to preserve rare books rather than destroy them: the volumes are kept in a library and scanned non-destructively, while the extracted text still flows into AI training data.

Both pipelines end in the same place — books converted into training tokens — but the difference matters. Non-destructive scanning is slower and costlier per volume, yet it leaves the physical artifact intact for the future. Anthropic optimized for throughput; Musk's team, at least where rare books are concerned, optimized for preservation. As AI labs digest ever-larger portions of the printed record, which of those instincts becomes the industry norm is more than an aesthetic question.

Why Books Are the AI Industry's Most Wanted Data

Beneath both stories sits the same pressure: frontier labs are running low on high-quality text. Books are the densest, cleanest source available — long-form, edited, and coherent in a way most web content is not. Rare and antique volumes add something the open web cannot: text that may never have been digitized anywhere, invisible to every competitor scraping the same internet.

That scarcity is why physical books — objects the technology industry spent two decades declaring obsolete — have quietly become a strategic asset. Whoever converts more of the printed record into training data first gets tokens nobody else has. Project Panama is what that race looks like in practice: buying the past in bulk and feeding it into a model.

![A digital library interface with a selection of books available for reading](https://supabase.srv1729373.hstgr.cloud/storage/v1/object/public/blog-images/speka-info/anthropic-project-panama-book-scanning-3-e1d59b9f83a9e0cf.png)

What It Means for Builders, Learners, and the Rest of Us

For everyday Claude users, nothing visibly changes — the text in question has already shaped the models people use today. But training-data provenance is no longer an academic concern. Developers who assemble products on top of frontier models, from API-level integrations to no-code agent platforms like [Hyperagent](https://speka.info/blog/hyperagent-no-code-ai-agent-platform-from-airtable-founder), inherit the data decisions of the labs underneath them — including the legal exposure and the ethical baggage those decisions carry.

The open-source world is part of this conversation too. Community model and dataset projects tend to document their training sources far more transparently than frontier labs do, and demand for that transparency keeps growing — a theme that surfaces repeatedly in the projects we covered in our roundup of [trending open-source AI repos](https://speka.info/blog/6-trending-open-source-ai-repos-on-github-this-week). And now that training-data stories are mainstream news, newcomers are encountering these questions from day one; anyone building that foundational understanding will find useful grounding in resources like [Microsoft's AI For Beginners course](https://speka.info/blog/microsoft-ai-for-beginners-course-tops-github-trending), which recently topped GitHub's trending charts.

The larger takeaway is that the era of quietly scraping whatever was reachable is ending. Between a $1.5 billion piracy payout and a book-buying operation industrial enough to earn its own codename, Anthropic has demonstrated both the cost of the old data playbook and the strange shape of the new one. We will keep tracking where the training-data race goes next — along with every major model release — in our [LLM Launches & Updates](https://speka.info/llm-updates/) coverage.

Frequently Asked Questions

What is Anthropic's Project Panama?

According to a Washington Post report, it was a secret Anthropic program that bought massive quantities of physical books, cut off their spines for high-speed scanning, recycled the destroyed copies, and used the digitized text to train Claude.

Why did Anthropic destroy the books instead of preserving them?

Cutting off a book's spine lets loose pages run through industrial sheet scanners far faster than scanning a bound volume. Destruction was a speed and cost decision, and the remains were recycled after digitization.

Why did Anthropic buy physical books instead of ebooks?

Ebooks come with licensing restrictions on copying and text extraction. Owning a physical copy sidesteps those limits, and Anthropic used the scanned text for training without redistributing the scans.

What was Anthropic's $1.5 billion payout about?

Before buying books, Anthropic downloaded more than 7 million pirated books for pre-training. The resulting copyright lawsuit ended in a $1.5 billion payout, which prompted the shift to purchasing physical copies.

How does Elon Musk's book-scanning approach differ from Anthropic's?

Musk reportedly instructed his team to preserve rare books in a library and scan them non-destructively while still using the text as AI training data — the opposite of Anthropic's destructive cut-and-recycle pipeline.

← Back to all posts