Anthropic’s book-scanning operation, known internally as Project Panama, exposes a starkly physical side of the generative AI race: millions of paper books bought in bulk, mechanically disbound, fed through industrial scanners, digitized for use in building Claude, and then recycled. The newly public details turn an abstract debate about AI training data into a concrete supply chain—one that begins in used-book warehouses and ends with searchable files, tokenized text, and discarded pages. Court filings and reporting reviewed by The Washington Post show that the project began in early 2024 after Anthropic confronted the legal exposure created by its earlier acquisition of pirated book collections.
That sequence matters. Project Panama was not merely a digitization project; it was a sourcing strategy shaped by copyright risk, data-quality demands, and the extraordinary scale required for large-language-model training. The result may offer an important legal and operational playbook for AI companies, but it also raises uncomfortable questions for readers, writers, libraries, collectors, and the used-book ecosystem.

An automated archive facility scans old books and documents with robotic machinery and AI data-provenance displays.Overview: The AI Training Data Problem Comes Off the Screen​

Large language models such as Anthropic’s Claude need immense quantities of text during pretraining. Books are particularly valuable because they contain long-form, edited, structured, and generally coherent writing—qualities that are difficult to reproduce through a simple web crawl. In the Bartz v. Anthropic litigation, the court record describes Anthropic’s view that books were a cost-effective route to a world-class model and that its customers wanted writing that was accurate, compelling, organized, and editor-approved. Judge William Alsup’s fair-use order records that internal rationale in unusual detail.
The technical process is more involved than “put books into an AI.” A physical volume first becomes a scanned image, then text through optical character recognition, then a cleaned dataset stripped of repetitive elements such as headers and page numbers. The text is converted into numerical tokens, which the model uses to learn statistical relationships among words and fragments of words during training. The court described multiple copying stages—from a central library, to working and cleaned copies, to tokenized versions used repeatedly in training. The underlying order makes clear that a book’s journey through an AI data pipeline is not a single act of copying.
For Windows users and IT professionals, this is a useful reminder that AI infrastructure is not just GPUs, cloud regions, and model weights. It also depends on data acquisition, metadata, rights management, storage, document processing, OCR quality, deduplication, provenance tracking, and secure deletion. Project Panama is an extreme example of that often-overlooked layer of the AI stack.

Project Panama: Buy, Cut, Scan, Recycle​

The most arresting detail in the reporting is the project’s destructive scanning model. Internal planning materials described Project Panama as an effort to “destructively scan all the books in the world,” while also directing employees not to publicize the effort. The Washington Post’s account of the unsealed filings reports that Anthropic sought large-scale physical-book acquisition and scanning capacity rather than simply licensing a conventional e-book collection.
The industrial workflow was blunt but efficient:
  1. Acquire used physical books in large quantities.
  2. Remove the spine using a hydraulic cutting machine.
  3. Separate pages for high-speed production scanning.
  4. Convert the scans into searchable digital files.
  5. Recycle the remaining paper.
A vendor proposal referenced an ability to convert between 500,000 and two million books over six months. That figure should be treated as a proposal-level estimate rather than a confirmed total for books ultimately processed, because the public documents reportedly redact the final count and spending details. Still, it gives a sense of the intended operational scale. The reporting on the vendor materials describes the cutting, scanning, and recycling sequence directly.

Why destruction was operationally attractive​

From a data-center and logistics perspective, destructive scanning has obvious advantages. Bound books are slow to digitize without damaging them. Non-destructive scanning requires careful handling, specialized overhead imaging, page-turning systems, manual intervention, and much more time per volume.
Cutting the spine turns a book into a stack of loose sheets. That enables high-throughput document scanners to process pages at industrial speeds with more consistent image quality and less human labor. It also eliminates the need to store millions of source books after digitization.
For a company trying to build a central training-data library, that is a powerful economic equation:
  • Used books can be cheaper than newly licensed digital editions.
  • Bulk purchases simplify transaction management.
  • Physical copies can be transformed into searchable records.
  • A destroyed source copy reduces warehousing costs.
  • Digital text is easier to filter, deduplicate, tokenize, and send through a machine-learning pipeline.
The system is efficient in a narrow operational sense. Yet efficiency is not the same as wisdom, cultural responsibility, or a complete answer to copyright concerns.

The booksellers and the scale effect​

Reporting indicates that Anthropic acquired books in batches often numbering in the tens of thousands, with used-book retailers including Better World Books and World of Books identified in the filings. The Washington Post’s reporting notes that the final number purchased and the amount spent were redacted.
This is precisely where the story extends beyond Anthropic. When a well-funded AI developer enters secondhand supply channels, it does not buy like an ordinary reader, student, local library, or rare-book collector. It can buy at volume, bid more aggressively, use intermediaries, and reduce available inventory quickly.
That does not automatically make every bulk purchase improper. A legally bought book is ordinarily a lawful object of resale. But it changes the market dynamic when the buyer’s goal is not reading, lending, preservation, or resale—it is extracting the text at scale and destroying the particular physical artifact after processing.

The Piracy Background: Why Anthropic Pivoted to Physical Books​

Project Panama did not emerge in a vacuum. It followed a period in which Anthropic acquired extensive collections from online “shadow libraries,” including Library Genesis (LibGen) and the Pirate Library Mirror (PiLiMi). Judge Alsup’s June 2025 order states that Anthropic downloaded at least five million book copies from LibGen in June 2021 and at least two million more from PiLiMi in July 2022, totaling more than seven million pirated copies in the court’s account. The order’s factual findings are explicit on both the sources and the scale.
The court’s record is especially consequential because it distinguishes between separate acts that can easily be blurred together in public debate:
  • Downloading and retaining unauthorized copies.
  • Maintaining a broad internal “central library.”
  • Copying particular works for training runs.
  • Training a model using lawfully sourced material.
  • Producing model outputs that might or might not reproduce protected expression.
Those distinctions drove the legal outcome. The court did not endorse a general proposition that any company may download whatever it wants from pirate sites so long as it calls the eventual use “AI research.” In fact, the ruling sharply rejected that approach. Judge Alsup’s order concluded that piracy of otherwise available books is not excused merely because the copies might later support a transformative use.

“All the books in the world” and the search for a lawful route​

The same ruling records that Anthropic hired Tom Turvey, previously the head of partnerships for Google’s book-scanning project, in February 2024. His task, according to the order, was to help obtain “all the books in the world” while avoiding as much “legal/practice/business slog” as possible. The court’s factual narrative places that hiring directly within Anthropic’s shift away from reliance on pirated collections.
That phrase captures the strategic tension at the heart of modern AI development. A frontier-model developer wants broad and diverse training data. Licensing millions of individual titles can be slow, expensive, administratively complicated, and limited by fragmented rights ownership. But bypassing those difficulties through unauthorized acquisition creates an enormous legal and reputational risk.
Project Panama appears to be Anthropic’s answer: purchase the physical object, digitize it internally, and use the resulting digital corpus rather than depending on unlicensed downloads.

What the Fair-Use Ruling Actually Said​

The June 2025 decision in Bartz v. Anthropic is frequently summarized as a sweeping victory for AI companies. That description is incomplete.
Judge Alsup held that using the books at issue to train Claude and its predecessors was “exceedingly transformative” and therefore a fair use in the circumstances before the court. The ruling reasoned that the model used books to generate something different rather than to reproduce or substitute for the books themselves. The fair-use order also noted that the authors had not alleged public outputs that were exact copies or infringing knockoffs of their books.
The court separately concluded that purchased print books digitized for Anthropic’s internal library were fair use. Its reasoning was notably practical: Anthropic had purchased the books, converting them into more convenient and searchable digital versions, copying the full work was necessary for that purpose, and the physical source copies were destroyed. The court’s analysis of purchased print-library copies treated the scan as a format shift rather than an unauthorized effort to create an additional competing library.
That distinction is the legal hinge of Project Panama.

The court’s limits were just as important​

The ruling did not find that Anthropic’s pirated central library was fair use. The court held that Anthropic lacked entitlement to retain those copies, particularly where they were obtained from pirate sources and kept even after the company had decided not to use them for model training. Judge Alsup’s discussion of the pirated library characterized the retention of a general-purpose collection as a separate, non-transformative use.
In plain English, the court drew a bright practical line:
  • Buying a physical book and scanning it for an internal digital collection was treated as legally permissible in this case.
  • Downloading a book from a pirate repository and retaining it as part of a broad library was not excused simply because AI training was somewhere in the chain of intended uses.
That is a narrower and more fact-specific conclusion than the slogan “AI training is fair use” suggests.

Why the ruling is influential—but not a universal license​

The decision has obvious value to AI companies. It offers a model for acquiring text without negotiating a license for every work: purchase books lawfully, digitize them, and keep the use internal. Yet it is not a statute, and it does not resolve every copyright question surrounding every model, dataset, or output.
The ruling depended on crucial facts, including the nature of the training use, the absence of alleged infringing outputs in the case, and the destruction of source copies in the purchased-print workflow. Future cases may focus on materially different facts, such as:
  • Model memorization and verbatim output.
  • Circumvention of technical protections.
  • Distribution of scanned source files.
  • Dataset resale or external access.
  • Market harm to licensing businesses.
  • Different classes of works, including images, newspapers, software, or unpublished materials.
The broader lesson for technology leaders is not “copyright no longer matters.” It is that provenance and workflow design now matter as much as model architecture.

The $1.5 Billion Settlement and Its Meaning​

Anthropic later agreed to a $1.5 billion settlement with authors and publishers over claims concerning pirated book files, without admitting wrongdoing. The settlement reporting described the agreement as resolving the remaining legacy claims tied to the unauthorized acquisition and storage allegations rather than overturning the fair-use holding on training.
The settlement structure illustrates the practical cost of treating copyright sourcing as an afterthought. Court reporting described a fund intended to allocate compensation based on eligible works, with an estimated payment in the neighborhood of several thousand dollars per covered title. Coverage of the approval process also reported that the deal required destruction of original files associated with the pirated datasets and certification concerning their use in commercially released models.
For an AI industry accustomed to discussing training data in terms of terabytes, tokens, and benchmarks, the settlement translates risk into a different unit: liability per work.

A warning against “data first, permissions later”​

The case reinforces a basic enterprise technology principle: obtaining the data lawfully at the beginning is less expensive than resolving provenance failures after a product has become strategically important.
Organizations building AI systems should treat training data with the same seriousness they apply to:
  • Software supply-chain security.
  • Open-source license compliance.
  • Personally identifiable information controls.
  • Records-retention policies.
  • Security logging and audit trails.
  • Vendor due diligence.
A dataset without reliable provenance can become a latent legal, financial, and reputational vulnerability. It may also become technically difficult to unwind after the data has been copied into preprocessing systems, fine-tuning jobs, evaluation datasets, retrieval systems, and model-development archives.

The Cultural and Market Risks of Destructive Scanning​

The legal question is only one part of the Project Panama story. The cultural question is harder: what is lost when a physical book is treated as raw material?
For ordinary mass-market paperbacks with millions of surviving copies, destruction may appear unremarkable. Libraries routinely cull duplicates, damaged items, and low-demand titles. Used-book sellers also recycle books they cannot profitably store or sell.
But the risk changes dramatically for out-of-print works, unusual editions, annotated copies, regional publishing, small-press titles, specialist technical manuals, and books whose rarity is not evident from an ISBN listing or a warehouse inventory record.

A scan is not the same as the book​

Digitization preserves text imperfectly; it does not preserve the complete physical object. A paper book can carry evidentiary and cultural value beyond its words:
  • Edition-specific typography and layout.
  • Dust jackets, cover art, and design.
  • Marginalia, inscriptions, bookplates, and stamps.
  • Original pagination and illustrations.
  • Paper stock, binding, and printing methods.
  • Publishing history and provenance.
  • Evidence of how a work was marketed, sold, circulated, and read.
An OCR-derived text file may be enough to support language-model training. It is not necessarily enough for bibliographers, conservators, historians, collectors, or future scholars.
That is why the phrase “destructively scan all the books in the world” lands differently from ordinary digitization. It suggests a model in which the physical archive is disposable once its linguistic content has been extracted.

The rare-book problem is not hypothetical in principle​

Used-book markets are fragmented. Sellers do not always know which books are scarce, historically important, or difficult to replace. Bulk acquisition magnifies that problem because it reduces the chance of title-by-title appraisal before processing.
The reporting does not establish that Anthropic knowingly destroyed one-of-a-kind volumes. It would be irresponsible to claim that without evidence. But the workflow plainly creates a preservation risk whenever high-volume buyers acquire books faster than they can be evaluated for rarity, condition, edition history, or local significance.
A responsible destructive-scanning program would therefore need safeguards well beyond normal procurement:
  1. Automated rarity and holdings checks against library catalogs and major bibliographic databases.
  2. Manual review thresholds for older, unusual, signed, annotated, or low-circulation titles.
  3. Exclusion lists for special collections, archival materials, and culturally sensitive works.
  4. Resale or donation channels for books unsuitable for destruction.
  5. Auditable chain-of-custody records for every acquired volume.
  6. Independent preservation oversight rather than relying solely on a buyer’s internal commercial incentives.
Without such controls, the pressure to optimize throughput can make irreversible loss look like a routine operational detail.

What Project Panama Means for AI, Windows, and Enterprise IT​

For WindowsForum readers, the immediate story is not that desktop PCs will suddenly become book scanners. The larger implication is that AI’s most contested component may increasingly be the data pipeline behind the model, not only the application interface in front of users.
Microsoft, Google, Anthropic, OpenAI, Meta, and other AI developers compete on model performance, reliability, safety, price, and integration. But each also competes for access to high-quality human-created material. The more the public web fills with machine-generated text, duplicated pages, SEO spam, and automated content farms, the greater the perceived value of curated human-authored sources such as books.
That can create a feedback loop:
  • AI systems are trained on high-quality human writing.
  • AI systems generate enormous quantities of new text.
  • The open web becomes noisier and harder to trust.
  • Publishers and curated archives become more valuable as training sources.
  • AI companies seek more direct, controlled, legally defensible acquisition channels.
Project Panama demonstrates one possible answer: turn the global used-book supply into an input stream for AI development.

Data governance becomes product strategy​

Enterprise IT teams should take note. The same questions that surround frontier-model training are moving downstream into corporate AI deployments:
  • Where did the documents in a retrieval-augmented generation system originate?
  • Does the organization have the right to digitize, index, summarize, and embed them?
  • Which files are retained after a pilot ends?
  • Can data be traced through preprocessing, fine-tuning, and evaluation stages?
  • Are employees uploading copyrighted manuals, books, reports, or proprietary customer documents into external AI tools?
The old distinction between “document management” and “AI development” is eroding. Every searchable corporate archive can become a potential model input, retrieval corpus, or fine-tuning dataset. That makes records management, licensing, and retention policy foundational parts of an AI strategy.

The Central Contradiction​

Anthropic’s destructive book scanning is easy to view as either a clever lawful workaround or an alarming act of cultural extraction. Both views contain elements of truth.
The strongest case for the practice is that lawful purchase matters. Buying books, scanning them internally, and avoiding redistribution is materially different from downloading millions of unauthorized files. The Bartz ruling recognized that distinction and treated the purchased-print conversion as fair use under the facts of the case. The court’s fair-use analysis gives AI developers a concrete incentive to invest in lawful acquisition rather than piracy.
The strongest case against it is that the legal right to scan a purchased copy does not settle the ethical, cultural, or market consequences of destroying the copy. A book is simultaneously a copyrighted work, a purchasable product, a physical artifact, a piece of publishing history, and sometimes a scarce cultural record. A process optimized for model training sees primarily the first two.
Project Panama therefore reveals the central contradiction of the AI era. The industry wants to build systems that can write, reason, summarize, and assist using the accumulated work of human authors. Yet the industrial process used to acquire that knowledge can treat the physical embodiments of that work as expendable.
The court ruling may have given companies a legal route around one part of the training-data problem. It did not answer the larger question of stewardship. As AI companies intensify their search for high-quality human text, the next test will be whether they can build lawful data pipelines without turning irreplaceable parts of the world’s printed record into recyclable feedstock.

References​

  1. Primary source: International Business Times UK
    Published: 2026-07-28T21:20:02+00:00
  2. Related coverage: washingtonpost.com
  3. Related coverage: prensa.com
  4. Related coverage: conven.org