Project Panama: How Anthropic secretly destroyed millions of books to train its AI

5 hours ago  ·  4 min read
By Jessica Johnson - usagevpn.com
1200x675_cmsv2_f549b144-3870-58ba-b31b-fee11d90a1c4-9863466

Millions of Books Met Their End in Anthropic’s Secret AI Training Operation

Usagevpn.com – Inside sprawling American warehouses, an unusual transformation was underway. Book spines were being sliced apart with mechanical precision, their pages separated and fed through high-speed scanners. What once stood as physical volumes on shelves became streams of digital data, destined to teach artificial intelligence how to write like humans. The paper, having served its purpose, was then sent to recycling facilities to become toilet paper and cardboard boxes.

This was Project Panama, Anthropic’s covert initiative to scan and destroy books worldwide. Internal documents revealed in late July legal filings exposed the company’s ambitious plan. The name itself suggested both a clandestine mission and financial transparency efforts. “We don’t want it to be known that we are working on this,” one planning document stated, capturing the secretive nature of the operation.

The Quest for Quality Training Data

Project Panama launched in early 2024 when Anthropic executives identified a critical gap. They needed books predating those already available online—materials that could teach Claude, their AI chatbot, to write with genuine quality rather than mimicking what they called “low quality internet speak.”

The company purchased tens of thousands of volumes simultaneously from various vendors, including Better World Books and the UK-based World of Books. A specialized vendor handled the physically demanding work of scanning and destroying the books. Hydraulic machines removed book spines while pages were sliced individually, then rapidly copied and uploaded using industrial imaging equipment.

One vendor working with Anthropic noted the company sought “an experienced document scanning services vendor to convert from 500,000 to two million books over a six-month period.” This massive undertaking reflected Anthropic’s confidence in the project’s potential.

Internal documents revealed executives believed this approach would improve communication quality and address “model collapse”—a degradation phenomenon occurring when AI systems train on text contaminated by their own generated output. Unlike much of the modern internet, books contain vast quantities of human-written, professionally edited literature largely untouched by AI-generated content.

Copyright Concerns and Legal Battles

The challenge for AI companies was acquiring this content. Court filings suggest Anthropic executives believed they could proceed without permission, and even if challenged, that the results would justify their methods.

One filing revealed Anthropic co-founder Ben Mann acquired copyright-protected material from a “shadow library” called LibGen in June 2021. The document included a screenshot showing Mann’s web browser downloading pirated books. The following year, Mann shared a link to the Pirate Library Mirror—a comprehensive catalogue of illegally obtained books—with fellow Anthropic employees, writing simply: “just in time!!!”

Project Panama only entered public awareness through thousands of pages of court documents unsealed in a copyright lawsuit filed by authors. In 2024, a group led by novelist Andrea Bartz and nonfiction writers Charles Graeber and Kirk Wallace Johnson filed a class action lawsuit against Anthropic. They alleged the company violated copyright laws by using pirated books to train its AI model.

Last month, Anthropic—currently valued at US$965 billion—agreed to pay $1.5 billion to settle the case, covering approximately 500,000 eligible works. “It is the largest known copyright recovery in history,” said Justin Nelson, the authors’ lead attorney.

A Broader Legal Landscape

This settlement represents only the beginning of a wave of litigation against AI companies. Meta, OpenAI, Google, and Microsoft all face similar copyright lawsuits from authors alleging piracy. Several authors who declined the class action settlement maintain individual cases against Anthropic.

In the original ruling last year, a judge determined that the company’s use of lawfully purchased books qualified as fair use under US law. However, the decision drew important distinctions between different categories of content and usage methods.

As additional cases progress through the judicial system, last week’s settlement may prove to be merely the opening chapter in what could become a prolonged legal struggle over the future of AI training data and the rights of creative professionals whose work powers these emerging technologies.

Frequently Asked Questions

What is Project Panama?

Project Panama is the main topic of this guide. The article explains the context, practical details, and next steps readers should understand.

Why does Project Panama matter?

Project Panama matters because readers are looking for a useful answer, not just a short summary. Good content should match search intent and help them decide what to do next.

MORE FROM THIS CATEGORY