Anthropic’s book-scanning operation, known internally as Project Panama, exposes a starkly physical side of the generative AI race: millions of paper books bought in bulk, mechanically disbound, fed through industrial scanners, digitized for use in building Claude, and then recycled. The newly public details turn an abstract debate about AI training data into a concrete supply chain—one that begins in used-book warehouses and ends with searchable files, tokenized text, and discarded pages. Court filings and reporting reviewed by The Washington Post show that the project began in early 2024 after Anthropic confronted the legal exposure created by its earlier acquisition of pirated book collections.
That sequence matters. Project Panama was not merely a digitization project; it was a sourcing strategy shaped by copyright risk, data-quality demands, and the extraordinary scale required for large-language-model training. The result may offer an important legal and operational playbook for AI companies, but it also raises uncomfortable questions for readers, writers, libraries, collectors, and the used-book ecosystem.
Large language models such as Anthropic’s Claude need immense quantities of text during pretraining. Books are particularly valuable because they contain long-form, edited, structured, and generally coherent writing—qualities that are difficult to reproduce through a simple web crawl. In the Bartz v. Anthropic litigation, the court record describes Anthropic’s view that books were a cost-effective route to a world-class model and that its customers wanted writing that was accurate, compelling, organized, and editor-approved. Judge William Alsup’s fair-use order records that internal rationale in unusual detail.
The technical process is more involved than “put books into an AI.” A physical volume first becomes a scanned image, then text through optical character recognition, then a cleaned dataset stripped of repetitive elements such as headers and page numbers. The text is converted into numerical tokens, which the model uses to learn statistical relationships among words and fragments of words during training. The court described multiple copying stages—from a central library, to working and cleaned copies, to tokenized versions used repeatedly in training. The underlying order makes clear that a book’s journey through an AI data pipeline is not a single act of copying.
For Windows users and IT professionals, this is a useful reminder that AI infrastructure is not just GPUs, cloud regions, and model weights. It also depends on data acquisition, metadata, rights management, storage, document processing, OCR quality, deduplication, provenance tracking, and secure deletion. Project Panama is an extreme example of that often-overlooked layer of the AI stack.
The industrial workflow was blunt but efficient:
Cutting the spine turns a book into a stack of loose sheets. That enables high-throughput document scanners to process pages at industrial speeds with more consistent image quality and less human labor. It also eliminates the need to store millions of source books after digitization.
For a company trying to build a central training-data library, that is a powerful economic equation:
This is precisely where the story extends beyond Anthropic. When a well-funded AI developer enters secondhand supply channels, it does not buy like an ordinary reader, student, local library, or rare-book collector. It can buy at volume, bid more aggressively, use intermediaries, and reduce available inventory quickly.
That does not automatically make every bulk purchase improper. A legally bought book is ordinarily a lawful object of resale. But it changes the market dynamic when the buyer’s goal is not reading, lending, preservation, or resale—it is extracting the text at scale and destroying the particular physical artifact after processing.
The court’s record is especially consequential because it distinguishes between separate acts that can easily be blurred together in public debate:
That phrase captures the strategic tension at the heart of modern AI development. A frontier-model developer wants broad and diverse training data. Licensing millions of individual titles can be slow, expensive, administratively complicated, and limited by fragmented rights ownership. But bypassing those difficulties through unauthorized acquisition creates an enormous legal and reputational risk.
Project Panama appears to be Anthropic’s answer: purchase the physical object, digitize it internally, and use the resulting digital corpus rather than depending on unlicensed downloads.
Judge Alsup held that using the books at issue to train Claude and its predecessors was “exceedingly transformative” and therefore a fair use in the circumstances before the court. The ruling reasoned that the model used books to generate something different rather than to reproduce or substitute for the books themselves. The fair-use order also noted that the authors had not alleged public outputs that were exact copies or infringing knockoffs of their books.
The court separately concluded that purchased print books digitized for Anthropic’s internal library were fair use. Its reasoning was notably practical: Anthropic had purchased the books, converting them into more convenient and searchable digital versions, copying the full work was necessary for that purpose, and the physical source copies were destroyed. The court’s analysis of purchased print-library copies treated the scan as a format shift rather than an unauthorized effort to create an additional competing library.
That distinction is the legal hinge of Project Panama.
In plain English, the court drew a bright practical line:
The ruling depended on crucial facts, including the nature of the training use, the absence of alleged infringing outputs in the case, and the destruction of source copies in the purchased-print workflow. Future cases may focus on materially different facts, such as:
The settlement structure illustrates the practical cost of treating copyright sourcing as an afterthought. Court reporting described a fund intended to allocate compensation based on eligible works, with an estimated payment in the neighborhood of several thousand dollars per covered title. Coverage of the approval process also reported that the deal required destruction of original files associated with the pirated datasets and certification concerning their use in commercially released models.
For an AI industry accustomed to discussing training data in terms of terabytes, tokens, and benchmarks, the settlement translates risk into a different unit: liability per work.
Organizations building AI systems should treat training data with the same seriousness they apply to:
For ordinary mass-market paperbacks with millions of surviving copies, destruction may appear unremarkable. Libraries routinely cull duplicates, damaged items, and low-demand titles. Used-book sellers also recycle books they cannot profitably store or sell.
But the risk changes dramatically for out-of-print works, unusual editions, annotated copies, regional publishing, small-press titles, specialist technical manuals, and books whose rarity is not evident from an ISBN listing or a warehouse inventory record.
That is why the phrase “destructively scan all the books in the world” lands differently from ordinary digitization. It suggests a model in which the physical archive is disposable once its linguistic content has been extracted.
The reporting does not establish that Anthropic knowingly destroyed one-of-a-kind volumes. It would be irresponsible to claim that without evidence. But the workflow plainly creates a preservation risk whenever high-volume buyers acquire books faster than they can be evaluated for rarity, condition, edition history, or local significance.
A responsible destructive-scanning program would therefore need safeguards well beyond normal procurement:
Microsoft, Google, Anthropic, OpenAI, Meta, and other AI developers compete on model performance, reliability, safety, price, and integration. But each also competes for access to high-quality human-created material. The more the public web fills with machine-generated text, duplicated pages, SEO spam, and automated content farms, the greater the perceived value of curated human-authored sources such as books.
That can create a feedback loop:
The strongest case for the practice is that lawful purchase matters. Buying books, scanning them internally, and avoiding redistribution is materially different from downloading millions of unauthorized files. The Bartz ruling recognized that distinction and treated the purchased-print conversion as fair use under the facts of the case. The court’s fair-use analysis gives AI developers a concrete incentive to invest in lawful acquisition rather than piracy.
The strongest case against it is that the legal right to scan a purchased copy does not settle the ethical, cultural, or market consequences of destroying the copy. A book is simultaneously a copyrighted work, a purchasable product, a physical artifact, a piece of publishing history, and sometimes a scarce cultural record. A process optimized for model training sees primarily the first two.
Project Panama therefore reveals the central contradiction of the AI era. The industry wants to build systems that can write, reason, summarize, and assist using the accumulated work of human authors. Yet the industrial process used to acquire that knowledge can treat the physical embodiments of that work as expendable.
The court ruling may have given companies a legal route around one part of the training-data problem. It did not answer the larger question of stewardship. As AI companies intensify their search for high-quality human text, the next test will be whether they can build lawful data pipelines without turning irreplaceable parts of the world’s printed record into recyclable feedstock.
That sequence matters. Project Panama was not merely a digitization project; it was a sourcing strategy shaped by copyright risk, data-quality demands, and the extraordinary scale required for large-language-model training. The result may offer an important legal and operational playbook for AI companies, but it also raises uncomfortable questions for readers, writers, libraries, collectors, and the used-book ecosystem.
Overview: The AI Training Data Problem Comes Off the Screen
Large language models such as Anthropic’s Claude need immense quantities of text during pretraining. Books are particularly valuable because they contain long-form, edited, structured, and generally coherent writing—qualities that are difficult to reproduce through a simple web crawl. In the Bartz v. Anthropic litigation, the court record describes Anthropic’s view that books were a cost-effective route to a world-class model and that its customers wanted writing that was accurate, compelling, organized, and editor-approved. Judge William Alsup’s fair-use order records that internal rationale in unusual detail.The technical process is more involved than “put books into an AI.” A physical volume first becomes a scanned image, then text through optical character recognition, then a cleaned dataset stripped of repetitive elements such as headers and page numbers. The text is converted into numerical tokens, which the model uses to learn statistical relationships among words and fragments of words during training. The court described multiple copying stages—from a central library, to working and cleaned copies, to tokenized versions used repeatedly in training. The underlying order makes clear that a book’s journey through an AI data pipeline is not a single act of copying.
For Windows users and IT professionals, this is a useful reminder that AI infrastructure is not just GPUs, cloud regions, and model weights. It also depends on data acquisition, metadata, rights management, storage, document processing, OCR quality, deduplication, provenance tracking, and secure deletion. Project Panama is an extreme example of that often-overlooked layer of the AI stack.
Project Panama: Buy, Cut, Scan, Recycle
The most arresting detail in the reporting is the project’s destructive scanning model. Internal planning materials described Project Panama as an effort to “destructively scan all the books in the world,” while also directing employees not to publicize the effort. The Washington Post’s account of the unsealed filings reports that Anthropic sought large-scale physical-book acquisition and scanning capacity rather than simply licensing a conventional e-book collection.The industrial workflow was blunt but efficient:
- Acquire used physical books in large quantities.
- Remove the spine using a hydraulic cutting machine.
- Separate pages for high-speed production scanning.
- Convert the scans into searchable digital files.
- Recycle the remaining paper.
Why destruction was operationally attractive
From a data-center and logistics perspective, destructive scanning has obvious advantages. Bound books are slow to digitize without damaging them. Non-destructive scanning requires careful handling, specialized overhead imaging, page-turning systems, manual intervention, and much more time per volume.Cutting the spine turns a book into a stack of loose sheets. That enables high-throughput document scanners to process pages at industrial speeds with more consistent image quality and less human labor. It also eliminates the need to store millions of source books after digitization.
For a company trying to build a central training-data library, that is a powerful economic equation:
- Used books can be cheaper than newly licensed digital editions.
- Bulk purchases simplify transaction management.
- Physical copies can be transformed into searchable records.
- A destroyed source copy reduces warehousing costs.
- Digital text is easier to filter, deduplicate, tokenize, and send through a machine-learning pipeline.
The booksellers and the scale effect
Reporting indicates that Anthropic acquired books in batches often numbering in the tens of thousands, with used-book retailers including Better World Books and World of Books identified in the filings. The Washington Post’s reporting notes that the final number purchased and the amount spent were redacted.This is precisely where the story extends beyond Anthropic. When a well-funded AI developer enters secondhand supply channels, it does not buy like an ordinary reader, student, local library, or rare-book collector. It can buy at volume, bid more aggressively, use intermediaries, and reduce available inventory quickly.
That does not automatically make every bulk purchase improper. A legally bought book is ordinarily a lawful object of resale. But it changes the market dynamic when the buyer’s goal is not reading, lending, preservation, or resale—it is extracting the text at scale and destroying the particular physical artifact after processing.
The Piracy Background: Why Anthropic Pivoted to Physical Books
Project Panama did not emerge in a vacuum. It followed a period in which Anthropic acquired extensive collections from online “shadow libraries,” including Library Genesis (LibGen) and the Pirate Library Mirror (PiLiMi). Judge Alsup’s June 2025 order states that Anthropic downloaded at least five million book copies from LibGen in June 2021 and at least two million more from PiLiMi in July 2022, totaling more than seven million pirated copies in the court’s account. The order’s factual findings are explicit on both the sources and the scale.The court’s record is especially consequential because it distinguishes between separate acts that can easily be blurred together in public debate:
- Downloading and retaining unauthorized copies.
- Maintaining a broad internal “central library.”
- Copying particular works for training runs.
- Training a model using lawfully sourced material.
- Producing model outputs that might or might not reproduce protected expression.
“All the books in the world” and the search for a lawful route
The same ruling records that Anthropic hired Tom Turvey, previously the head of partnerships for Google’s book-scanning project, in February 2024. His task, according to the order, was to help obtain “all the books in the world” while avoiding as much “legal/practice/business slog” as possible. The court’s factual narrative places that hiring directly within Anthropic’s shift away from reliance on pirated collections.That phrase captures the strategic tension at the heart of modern AI development. A frontier-model developer wants broad and diverse training data. Licensing millions of individual titles can be slow, expensive, administratively complicated, and limited by fragmented rights ownership. But bypassing those difficulties through unauthorized acquisition creates an enormous legal and reputational risk.
Project Panama appears to be Anthropic’s answer: purchase the physical object, digitize it internally, and use the resulting digital corpus rather than depending on unlicensed downloads.
What the Fair-Use Ruling Actually Said
The June 2025 decision in Bartz v. Anthropic is frequently summarized as a sweeping victory for AI companies. That description is incomplete.Judge Alsup held that using the books at issue to train Claude and its predecessors was “exceedingly transformative” and therefore a fair use in the circumstances before the court. The ruling reasoned that the model used books to generate something different rather than to reproduce or substitute for the books themselves. The fair-use order also noted that the authors had not alleged public outputs that were exact copies or infringing knockoffs of their books.
The court separately concluded that purchased print books digitized for Anthropic’s internal library were fair use. Its reasoning was notably practical: Anthropic had purchased the books, converting them into more convenient and searchable digital versions, copying the full work was necessary for that purpose, and the physical source copies were destroyed. The court’s analysis of purchased print-library copies treated the scan as a format shift rather than an unauthorized effort to create an additional competing library.
That distinction is the legal hinge of Project Panama.
The court’s limits were just as important
The ruling did not find that Anthropic’s pirated central library was fair use. The court held that Anthropic lacked entitlement to retain those copies, particularly where they were obtained from pirate sources and kept even after the company had decided not to use them for model training. Judge Alsup’s discussion of the pirated library characterized the retention of a general-purpose collection as a separate, non-transformative use.In plain English, the court drew a bright practical line:
- Buying a physical book and scanning it for an internal digital collection was treated as legally permissible in this case.
- Downloading a book from a pirate repository and retaining it as part of a broad library was not excused simply because AI training was somewhere in the chain of intended uses.
Why the ruling is influential—but not a universal license
The decision has obvious value to AI companies. It offers a model for acquiring text without negotiating a license for every work: purchase books lawfully, digitize them, and keep the use internal. Yet it is not a statute, and it does not resolve every copyright question surrounding every model, dataset, or output.The ruling depended on crucial facts, including the nature of the training use, the absence of alleged infringing outputs in the case, and the destruction of source copies in the purchased-print workflow. Future cases may focus on materially different facts, such as:
- Model memorization and verbatim output.
- Circumvention of technical protections.
- Distribution of scanned source files.
- Dataset resale or external access.
- Market harm to licensing businesses.
- Different classes of works, including images, newspapers, software, or unpublished materials.
The $1.5 Billion Settlement and Its Meaning
Anthropic later agreed to a $1.5 billion settlement with authors and publishers over claims concerning pirated book files, without admitting wrongdoing. The settlement reporting described the agreement as resolving the remaining legacy claims tied to the unauthorized acquisition and storage allegations rather than overturning the fair-use holding on training.The settlement structure illustrates the practical cost of treating copyright sourcing as an afterthought. Court reporting described a fund intended to allocate compensation based on eligible works, with an estimated payment in the neighborhood of several thousand dollars per covered title. Coverage of the approval process also reported that the deal required destruction of original files associated with the pirated datasets and certification concerning their use in commercially released models.
For an AI industry accustomed to discussing training data in terms of terabytes, tokens, and benchmarks, the settlement translates risk into a different unit: liability per work.
A warning against “data first, permissions later”
The case reinforces a basic enterprise technology principle: obtaining the data lawfully at the beginning is less expensive than resolving provenance failures after a product has become strategically important.Organizations building AI systems should treat training data with the same seriousness they apply to:
- Software supply-chain security.
- Open-source license compliance.
- Personally identifiable information controls.
- Records-retention policies.
- Security logging and audit trails.
- Vendor due diligence.
The Cultural and Market Risks of Destructive Scanning
The legal question is only one part of the Project Panama story. The cultural question is harder: what is lost when a physical book is treated as raw material?For ordinary mass-market paperbacks with millions of surviving copies, destruction may appear unremarkable. Libraries routinely cull duplicates, damaged items, and low-demand titles. Used-book sellers also recycle books they cannot profitably store or sell.
But the risk changes dramatically for out-of-print works, unusual editions, annotated copies, regional publishing, small-press titles, specialist technical manuals, and books whose rarity is not evident from an ISBN listing or a warehouse inventory record.
A scan is not the same as the book
Digitization preserves text imperfectly; it does not preserve the complete physical object. A paper book can carry evidentiary and cultural value beyond its words:- Edition-specific typography and layout.
- Dust jackets, cover art, and design.
- Marginalia, inscriptions, bookplates, and stamps.
- Original pagination and illustrations.
- Paper stock, binding, and printing methods.
- Publishing history and provenance.
- Evidence of how a work was marketed, sold, circulated, and read.
That is why the phrase “destructively scan all the books in the world” lands differently from ordinary digitization. It suggests a model in which the physical archive is disposable once its linguistic content has been extracted.
The rare-book problem is not hypothetical in principle
Used-book markets are fragmented. Sellers do not always know which books are scarce, historically important, or difficult to replace. Bulk acquisition magnifies that problem because it reduces the chance of title-by-title appraisal before processing.The reporting does not establish that Anthropic knowingly destroyed one-of-a-kind volumes. It would be irresponsible to claim that without evidence. But the workflow plainly creates a preservation risk whenever high-volume buyers acquire books faster than they can be evaluated for rarity, condition, edition history, or local significance.
A responsible destructive-scanning program would therefore need safeguards well beyond normal procurement:
- Automated rarity and holdings checks against library catalogs and major bibliographic databases.
- Manual review thresholds for older, unusual, signed, annotated, or low-circulation titles.
- Exclusion lists for special collections, archival materials, and culturally sensitive works.
- Resale or donation channels for books unsuitable for destruction.
- Auditable chain-of-custody records for every acquired volume.
- Independent preservation oversight rather than relying solely on a buyer’s internal commercial incentives.
What Project Panama Means for AI, Windows, and Enterprise IT
For WindowsForum readers, the immediate story is not that desktop PCs will suddenly become book scanners. The larger implication is that AI’s most contested component may increasingly be the data pipeline behind the model, not only the application interface in front of users.Microsoft, Google, Anthropic, OpenAI, Meta, and other AI developers compete on model performance, reliability, safety, price, and integration. But each also competes for access to high-quality human-created material. The more the public web fills with machine-generated text, duplicated pages, SEO spam, and automated content farms, the greater the perceived value of curated human-authored sources such as books.
That can create a feedback loop:
- AI systems are trained on high-quality human writing.
- AI systems generate enormous quantities of new text.
- The open web becomes noisier and harder to trust.
- Publishers and curated archives become more valuable as training sources.
- AI companies seek more direct, controlled, legally defensible acquisition channels.
Data governance becomes product strategy
Enterprise IT teams should take note. The same questions that surround frontier-model training are moving downstream into corporate AI deployments:- Where did the documents in a retrieval-augmented generation system originate?
- Does the organization have the right to digitize, index, summarize, and embed them?
- Which files are retained after a pilot ends?
- Can data be traced through preprocessing, fine-tuning, and evaluation stages?
- Are employees uploading copyrighted manuals, books, reports, or proprietary customer documents into external AI tools?
The Central Contradiction
Anthropic’s destructive book scanning is easy to view as either a clever lawful workaround or an alarming act of cultural extraction. Both views contain elements of truth.The strongest case for the practice is that lawful purchase matters. Buying books, scanning them internally, and avoiding redistribution is materially different from downloading millions of unauthorized files. The Bartz ruling recognized that distinction and treated the purchased-print conversion as fair use under the facts of the case. The court’s fair-use analysis gives AI developers a concrete incentive to invest in lawful acquisition rather than piracy.
The strongest case against it is that the legal right to scan a purchased copy does not settle the ethical, cultural, or market consequences of destroying the copy. A book is simultaneously a copyrighted work, a purchasable product, a physical artifact, a piece of publishing history, and sometimes a scarce cultural record. A process optimized for model training sees primarily the first two.
Project Panama therefore reveals the central contradiction of the AI era. The industry wants to build systems that can write, reason, summarize, and assist using the accumulated work of human authors. Yet the industrial process used to acquire that knowledge can treat the physical embodiments of that work as expendable.
The court ruling may have given companies a legal route around one part of the training-data problem. It did not answer the larger question of stewardship. As AI companies intensify their search for high-quality human text, the next test will be whether they can build lawful data pipelines without turning irreplaceable parts of the world’s printed record into recyclable feedstock.
References
- Primary source: International Business Times UK
Published: 2026-07-28T21:20:02+00:00
Inside Project Panama, Anthropic's Secret Effort To Scan and Shred the World's Books | IBTimes UK
Anthropic secretly scanned and destroyed millions of books for AI training, raising legal and ethical questions. The operation, Project Panama, involved both piracy and legal purchases, impactingwww.ibtimes.co.uk - Related coverage: washingtonpost.com
- Related coverage: prensa.com
‘Project Panama’: la operación secreta de Anthropic para comprar, escanear y destruir libros | La Prensa Panamá
Una investigación de The Washington Post, basada en documentos judiciales, revela cómo la empresa de inteligencia artificial Anthropic desarrolló una iniciativa denominada Project Panama para adquirir, desmembrar y di...www.prensa.com - Related coverage: conven.org