AI Companies Are Destroying Books. We Read the Court File.
By Milad Ghobadibeygvand, BScN (Western University, 2014) · Published September 6, 2026 · Zeus Media — Civic Duty

This investigation was researched and drafted with Claude, an AI system made by Anthropic — one of the two companies examined below. We are publishing the conflict instead of hiding it, and every claim in this piece traces to a court record, a produced document, a named report, or a count you can re-run yourself. Corporate statements appear here as exhibits of what companies chose to say, never as explanations we adopt. If any named party disputes a fact in this piece, write to milad@zeusebikes.ca and we will review it against the record — corrections are published, not buried.
In August 2026, the journalists at 404 Media hid a tracking device inside a shipment of rare books and followed it to an Amazon building outside Las Vegas whose team logo is a dinosaur holding a book. Seven months earlier, four thousand pages of court filings had been unsealed in San Francisco, and inside them sat a one-page company document that begins: "Project Panama is our effort to destructively scan all the books in the world."
If you read about any of this on your feed, you got the rumour version — shredders, secrecy, something about lawsuits. What you likely did not get is the record: who signed off, what the documents actually say, how many books are actually gone, whether any of it was legal, and whether your book — or your country's books — went into the pile. Deciding what to believe from a feed instead of a file is how every panic and every whitewash gets built.
So we pulled the file. All of it.
This piece is built on primary records and original counts, in that order: the full text of Judge William Alsup's fair-use order in Bartz v. Anthropic (Doc. 231, June 23, 2025); the exhibits unsealed on January 21, 2026 — including the Project Panama document itself (Bates ANT_BARTZ_000005492), the January 2023 executive email thread (Bates -000002604), vendor correspondence (Bates -035587–91), Tom Turvey's sworn declaration and 14 pages of his deposition; the public docket and its RECAP mirror; and the reporting of 404 Media, the Washington Post, Snopes, Fortune and the Las Vegas Review-Journal, each cited where used.
The numbers are ours where we could count them: we downloaded the February 12, 2021 LibGen metadata catalogue — bibliographic facts, not books; the archive itself contains no book text — and counted 2,911,058 non-fiction and 2,310,071 fiction records with an escape-aware parser, validating both totals against the Authors Alliance's independent tallies and the court's own figure. The Books3 filename roster (196,611 entries) sits in our corpus. Counting scripts carry provenance headers and are reproducible; on September 6 both counts were re-run on hashed copies of the extracted catalogues and reproduced to the record. Every file the piece relies on — thirty-two at publication, twenty-three of them external records — is listed with its route, retrieval date and SHA-256 hash in a manifest we hold; the chain-of-custody ledger below summarises the items the piece relies on, and every quotation from a scanned exhibit was checked against the page image, not against machine text. The settlement arithmetic is computed from the final-approval order of July 20, 2026, with every formula shown.
Three rules governed the writing. First: found, not found, and could-not-look are reported as three different things, and the could-not-look table is published below. Second: no motive is asserted anywhere in this piece unless a document states it — where the record shows only an action, we report the action and its date. Third: official statements are quoted as artifacts and tested against the record, never used as reasoning. Where we editorialize, the passage says so.
Yes — it is documented. Anthropic bought millions of print books, cut off their bindings, scanned them and discarded the paper under an internal program named Project Panama ("we don't want it to be known that we are working on this"), after pirating over seven million books first. A court blessed the buy-scan-destroy step and priced the piracy at $1.5 billion (US). Amazon runs a warehouse doing the same cut-scan-discard work; it has never named a reason or an executive. A title-level list of what died exists inside Anthropic, on the sworn record; a worker's account suggests one at Amazon too. Neither company has published a single title.
In this investigation
- Is it actually true?
- The document with the codename
- The names in the record
- How many books? The counts, verified
- Chain of custody: the evidence ledger
- The machine: vendors, cutters, Illinois
- The licensing road, opened and dropped
- Follow the money: the forensic ledger
- Amazon's warehouse and the accountability gap
- It is not two companies. It is a market.
- Is any of this legal?
- Was your book taken? How to check
- The lists exist — and no one will publish them
- The Canadian shelf
- Has this happened before? 2,200 years of destroyed books
- What the record does not show
- Frequently asked questions
- The bottom line — and the demand
Is It Actually True?
Yes. Two separate programs are documented — one in a federal court record, one by investigative reporting Amazon has not rebutted. Anthropic's program bought millions of print books, destroyed them in the scanning, and kept the digital copies; Amazon's facility receives bulk book shipments and cuts, scans and discards them. The rumour on your feed was, for once, smaller than the record.
Start with what a judge — not a journalist, not a company — put in writing. In June 2025, Senior US District Judge William Alsup summarized the undisputed facts of Bartz v. Anthropic:
"Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals."
— Order on Fair Use, Doc. 231 at 4, Bartz v. Anthropic, N.D. Cal., June 23, 2025
Both sides agreed on those facts — the fight was over what they meant legally. And on the Amazon side, the physical evidence is a tracking dot: 404 Media followed a shipment of rare books to the VGT3 facility, then interviewed workers inside it who described the machine that "just comes down and slices" the bindings, the 20-odd scanners, and the loose pages dumped into cardboard shuttles. Snopes, examining the rumour, reports that Amazon had previously denied engaging in destructive scanning; its current statement does not deny it.
Book destruction by AI companies is not an internet rumour. It is (a) an agreed fact in a US federal court record for Anthropic, and (b) tracked, witnessed and unrebutted reporting for Amazon. Everything else in this piece is detail on top of that foundation — names, counts, mechanics, law, and history.
The Document With the Codename
The single most important artifact in this story is one page long. It was written inside Anthropic, stamped "ATTORNEY CLIENT PRIVILEGED WORK PRODUCT," produced in discovery as Exhibit 21, sealed in April 2025, and refiled on the public docket on January 21, 2026. It reads, in full relevant part:
Project Panama
Owner: Tom Turvey
Last major update: Apr 13, 2024What is Project Panama?
Project Panama is our effort to destructively scan all the books in the world.Why use a codename?
We use a "soft codename" for it because we don't want it to be known that we are working on this. This document is visible to all Anthropic employees, but you should avoid talking about it in public areas, and the fact that we are working on this should not be shared with anyone outside Anthropic.— Doc. 554-21, Bates ANT_BARTZ_000005492, unsealed Jan. 21, 2026
Read it twice, because its structure tells you things no press release ever will. It names an owner. It has a version date. It links — in the original — to four internal trackers: a pipeline document with "token forecasts" and "batch operations info," a token tracker, a cost tracker (go/panama-costs), and a relative-value tracker. This was not a rogue experiment. It was an administered industrial program with accounting.
And it states, in the company's own grammar, that concealment was a design requirement. We do not have to characterize that. The sentence characterizes itself.
The Names in the Record
Who decided? The court file answers with named individuals — quoted below at exactly the strength of the source, with every softener kept in. Where the record shows an action, we report the action. Where it shows only silence, we report the silence. Operational staff who executed instructions are deliberately not named here; decision-makers are.
| Name | Role | What the record shows | Source |
|---|---|---|---|
| Dario Amodei | Co-founder & CEO, Anthropic | In a January 8, 2023 comment thread on "Anthropic Plan for 2023," responding to a colleague asking for a book-buying strategy: "I think we have many places from which we could in principle buy these things, but it's a huge legal/practice/business slog." He then marked the thread resolved. Judge Alsup, citing this exhibit, wrote that Anthropic "preferred to steal them" — the word "steal" is the judge's, the "slog" is Amodei's. | Doc. 554-27; Order at 2–3 |
| Ben Mann | Co-founder, Anthropic | Personally downloaded Books3 — 196,640 books he "knew had been assembled from unauthorized copies" — in January or February 2021, and "at least five million copies of books" from LibGen in June 2021, "which he knew had been pirated." Testified his fair-use belief was formed earlier, at OpenAI, where he was a GPT-3 architect. | Order at 2–3; deposition per case filings |
| Jared Kaplan | Co-founder & chief science officer, Anthropic | In the same January 2023 thread, named the legitimate routes then under negotiation: "Two places that sound promising are (1) buying [redacted] from ProQuest — we're negotiating this now and (2) Springer all access… But overall this is a hodgepodge." | Doc. 554-27 |
| Neerav Kingsland | Anthropic (role per company papers) | Opened that thread: "buy books and scientific papers. I think we need to have more specific approach here?… I don't know what our strategy is…" — and, after Amodei's "slog" reply: "Agree. Just naming I still have FUD here." | Doc. 554-27 |
| Tom Turvey | VP of Product Partnerships, Anthropic (from Feb. 2024); listed owner of the Project Panama document | Tasked, per the court, with getting "all the books in the world" while avoiding "legal/practice/business slog." Opened licensing talks with major publishers in spring 2024, then "let those conversations wither." Under oath: "We're not doing any book licensing directly with publishers at this point since we shifted our focus at Project Panama." | Order at 3–4; Depo, Doc. 554-22 |
| Amazon — no name | — | No court has compelled Amazon's book program into the record, and no executive has been named in any reporting. Its AI organization was led by SVP Rohit Prasad until his departure at the end of 2025, with Peter DeSantis leading the reorganized group since — stated here as organizational structure only, not as sourced involvement in the book program. | CNBC, Dec. 17, 2025 |
One asymmetry in that table deserves its own sentence: Anthropic's names are known because a federal court forced the record open; Amazon's are unknown because no court has. That is not a difference in corporate character. It is a difference in subpoena exposure — and it is measurable in this very article, where one company's section has quotations and the other's has a silence inventory.
The Turvey entry carries the story's sharpest fact, and it needs no adjective. His own sworn declaration lays out his résumé: eleven years in publishing, then twenty years at Google, where, in his words, he "helped create Google Books" — the 2004 project that scanned over 40 million books and did it non-destructively, with custom camera rigs, fragile volumes set aside, and every book returned to its library. His teams "concluded thousands of licensing deals" with publishers. The man who ran partnerships for the gentlest mass-digitization project in history is the listed owner of the document that begins "destructively scan all the books in the world." Both facts are his own account. We print the sequence; the reader can do the rest.
The company's own sworn paperwork adds a list of its own. Asked to identify the five people "most knowledgeable" about the 2024 book-scanning project, "including how You selected books for purchase/scanning," Anthropic answered on December 23, 2024 with five names — Kaplan first, Turvey third, and three operations and technical staff whom, consistent with the names policy above, we do not print (Doc. 554-30, Interrogatory 13). That answer gives Turvey's position as "Manager, Member of Technical Staff." His own declaration, sworn three months later, says he has been "Vice President of Product Partnerships" since February 2024. Two sworn documents, two titles; we print both and draw nothing from the gap.
October 2022: an Anthropic researcher writes in-channel that LibGen and PiLiMi are being dropped from the next dataset "for legal reasons." January 2023: the CEO calls lawful buying "a huge legal/practice/business slog," and the thread is resolved. The pirated libraries stayed on the servers — the court: "It kept them anyway." February 2024: Turvey is hired. April 2024: the Panama document is updated, secrecy clause included. The awareness always preceded the escalation.
How Many Books? The Counts, Verified
Over seven million pirated copies, plus millions of purchased-and-destroyed print books — and on the pirated side, we did not take anyone's word for it. Every number below is either the court's own figure or a count we ran ourselves on the public catalogues, and where both exist, they agree.
| Tranche | Court record | Independent check | Result |
|---|---|---|---|
| Books3 (downloaded Jan–Feb 2021) | 196,640 books (Order at 3) | Public filename roster on GitHub, counted by us | 196,611 filenames — delta of 29 against the court's count |
| LibGen (downloaded June 2021) | "at least five million copies" (Order at 3) | We counted the Feb. 12, 2021 catalogue snapshot: 2,911,058 non-fiction + 2,310,071 fiction records | 5,221,129 records — consistent with the court and with the Authors Alliance's ~2.9M/~2.3M tallies |
| PiLiMi (downloaded July 2022) | "at least two million copies" (Order at 3) | Not independently counted this round | Court figure stands alone |
| Pirated total | "over seven million copies of books" (Order at 3) | — | Three-way convergence on the LibGen leg |
| Print bought & destroyed | "millions of print books," "many millions of dollars" (Order at 4) | Washington Post, from 4,000+ unsealed pages: tens of millions of dollars; a vendor proposal to convert 500,000–2,000,000 books in six months | Unquantified beyond "millions" — the precise number sits in Anthropic's internal catalogue |
| Amazon | No court record exists | 404 Media: bulk shipments, ~20–25 scanners running; volume undisclosed | Unknown — and undisclosed is a choice |
A word on method, because it is the difference between journalism and vibes. The LibGen count was not scraped from someone's summary. We downloaded the February 2021 metadata snapshot — the catalogue vintage closest to the June 2021 download the court describes — from the Internet Archive's preserved copy, parsed 4.3 GB of database dumps with an escape-aware counter (a naive count overstates fiction by 2.4× because the dump bundles three tables — we caught that), and cross-checked against the Authors Alliance's independent analysis. Catalogue metadata is facts about books — titles, authors, ISBNs — not the books themselves; the archive item contains no book text, and neither does our corpus.
One honest limit, stated plainly: the catalogue tells you what LibGen held four months before the download — an upper-bound roster. The court says "at least five million copies" were taken; it does not say the entire collection was. The convergence of the numbers is strong evidence they are the same population. It is not identity, and we will not pretend otherwise.
Chain of Custody: The Evidence Ledger
Every document and dataset behind this piece is held in a single research directory, hashed, dated, and traced to the address it was retrieved from — thirty-two files as of publication, twenty-three of them external records and the rest our own scripts and outputs. This section summarises that manifest, because a number without a chain of custody is a rumour with a decimal point. Each row says what the item is, how we got it, what we did to check it, and what it can and cannot prove.
| Item | What it is | Route and date | Integrity check | Can prove / cannot prove |
|---|---|---|---|---|
| Doc. 231 | Judge Alsup's fair-use order, June 23, 2025, full text | CourtListener storage, Sept. 5, 2026 | SHA-256 recorded; every quotation checked against the text | Proves the agreed facts and the court's counts. Cannot prove the exact print total ("millions"). |
| Doc. 554-21 | The Project Panama document (Bates ANT_BARTZ_000005492) | RECAP mirror on the Internet Archive, Sept. 5, 2026 | Image PDF; OCR at 300 dpi, then every quoted line re-read against the page image; sealed-to-unsealed mapping confirmed via the April 2025 slip-sheet (Doc. 153-N → 554-N) | Proves the codename, the owner, the secrecy clause, the tracker links. Cannot prove what the trackers contained. |
| Doc. 554-27 | "Anthropic Plan for 2023" comment thread, Jan. 8, 2023 | RECAP mirror, Sept. 5 | OCR + image re-read; the notification email header carries the date | Proves the "slog" sentence and its date. Cannot prove intent beyond the words. |
| Doc. 554-22 | Turvey deposition excerpts, 14 pages, taken April 10, 2025 | RECAP mirror, Sept. 5 | OCR + image re-read; page numbers cited to the transcript's own pagination; redactions noted where the figure sits | Proves vendors, mechanics, access, the licensing pivot and the "more cost effective" answer. Cannot prove the per-book valuation ceiling — redacted. |
| Doc. 554-30 | Anthropic's sworn interrogatory responses, Dec. 23, 2024 | RECAP mirror, Sept. 5; first read Sept. 6 | Text PDF; quoted verbatim | Proves Anthropic agreed to produce a title-level list of scanned books used in training, that each book is scanned once, and who it names as most knowledgeable. Cannot prove the list's contents — produced under protective order. |
| Docs. 555-1, 555-2 | Vendor emails: Wonder Book (Aug. 2024); Editorial Océano de México (Aug. 2024) | RECAP mirror / CourtListener storage, Sept. 5 | OCR + image re-read; redaction bars noted | Proves what vendors proposed and what Anthropic asked for. Cannot prove volumes or prices — redacted. |
| Doc. 555-6 | ETH Zurich memorization disclosure chain, March 11–13, 2024 | CourtListener storage, Sept. 5; first read Sept. 6 | OCR + image re-read | Proves what the researchers reported and that the legal department said it was aware of the report. Cannot prove the researchers' findings — unadjudicated. |
| Doc. 127 | Turvey's sworn declaration, March 27, 2025 | CourtListener storage, Sept. 5 | Text PDF; paragraphs cited | Proves his career account and his "no viable market" defence. Cannot prove the purchase program's figures — ¶¶19–27 redacted. |
| Doc. 680 | Order granting final approval, fees and judgment, July 20, 2026, 23 pages | RECAP mirror, Sept. 6 | Text PDF; every dollar figure in the money section below is cited to its page | Proves the settlement arithmetic. Cannot prove when money moves — see the appeals. |
| Docs. 682, 683 | Notices of appeal to the Ninth Circuit, Aug. 18 and 19, 2026 | RECAP mirror, Sept. 6 | Text PDF | Proves that appeals focused on the fee order were filed. Cannot prove the effect on distribution timing. |
| Books3 roster | 196,611 filenames, compiled by a third party from the dataset | GitHub (psmedia/Books3Info), Sept. 5 | Line count re-run; three counts exist and all three are stated: our 196,611, the court's 196,640, the compiler's "197,500 txt files" | Proves which filenames were in the tranche. Cannot prove ISBN-level identity — filenames are "a little chaotic," the compiler says. |
| LibGen catalogue | Two metadata dumps, snapshot Feb. 12, 2021 — 1.25 GB of archives, 4.37 GB extracted; bibliographic records only, no book text | Internet Archive, Sept. 5; server confirmed complete transfer (HTTP 416 on resume) | Archives and extracted SQL both hashed; counted with an escape-aware parser; re-run Sept. 6 on the hashed files: 2,911,058 and 2,310,071, identical to the first run; validated against the Authors Alliance's independent tallies | Proves what LibGen held four months before the June 2021 download. Cannot prove that every record was downloaded. |
| Canadian count | 9,246-record floor, script and 400-row sample file | Our run, Sept. 5–6 | Denominators and field-population rates printed with the result; 20 of 20 sampled rows read by eye; 644 malformed rows (0.012%) skipped and declared | Proves a floor. Cannot prove a total — 83% of non-fiction rows have no city, fiction has no city field. |
| Washington Post | Jan. 27, 2026 investigation | Internet Archive snapshot of Aug. 21, 2026, Sept. 6 | Only the first two paragraphs survive in the snapshot's page data; both saved | Proves the Post's "tens of millions of dollars" line directly. The "500,000 to two million" vendor proposal and the "4,000 pages" figure reach us only through Euronews and IBTimes UK recaps of the Post — labelled as such wherever used. |
Four disciplines governed the handling. Found, not found, and could-not-look are three different states — a fetch that failed never became a zero. Image scans were never quoted from OCR alone; the machine text was a finding aid, and the page image was the source. Counts were validated against an outside denominator before use — the Authors Alliance's independent LibGen tallies, the court's Books3 figure — and one parser defect was caught this way: a naive count of the fiction dump reported 5.57 million because the file bundles three tables, an over-count of 2.4 times that a per-table counter corrected. Sealed-to-unsealed mapping was traced document by document, so that an exhibit cited as 554-N can be shown to be the same item filed under seal as 153-N in April 2025.
One discrepancy the custody work surfaced is printed in the names section above at exactly its strength: two sworn documents, three months apart, give Tom Turvey two different titles. Custody is how such things are found — by reading every page rather than the pages that were quoted elsewhere.
Thirty-two files in a hashed, dated manifest with a route for each. Two original counts re-run on hashed inputs and reproduced to the record. Every quotation from a scanned exhibit checked against the page image. If any number in this piece is wrong, the ledger says exactly which file to open to prove it.
The Machine: Vendors, Cutters, Illinois
The pipeline had names, addresses and shipping manifests, and Turvey described it under oath: books bought in bulk from three main suppliers, trucked to a scanning contractor in Illinois, cut apart by purpose-built German machines, photographed page by page, and thrown away. What follows is the supply chain of a book-destruction program, reconstructed entirely from sworn testimony and produced emails.

Who sold the books
Asked to list every English-language book company Project Panama acquired from, Turvey answered: "Baker & Taylor, World of Books, Better World. I think that's about it." Baker & Taylor — one of the oldest book distributors in North America — from "the summer of 2024," under an agreement for a volume that stays redacted; World of Books, a UK used-book reseller, from around November–December 2024; Better World Books, the used-book seller famous for its library partnerships, confirmed in staff emails ("We got [redacted] from BWB"). Barnes & Noble: no. ThriftBooks: no. Half Price Books was investigated and skipped — "the pricing was higher than… we were willing to pay."
The buying rule, sworn: "We would ask every vendor… please don't send any that are self-published or children's books or books that we have that would be redundant, like Bibles and the like." Vendors sent inventory lists; Anthropic deduplicated against what it already owned and took the remainder — when the examiner put it to him that "you would essentially just take the rest," Turvey agreed: "Yeah. Roughly."
The buying was not English-only, and the vendors' own numbers survive where Anthropic's are redacted. On August 6, 2024, an Anthropic product-partnerships staffer wrote to Editorial Océano de México, in Spanish and English, that "we're building a research library in a few languages" and wanted "a substantial volume of books, primarily in Spanish, to be shipped to the United States"; Océano's export manager replied the next morning with a stock catalogue and the offer of a volume discount (Doc. 555-2). Wonder Book, for its part, told Turvey its inventory was "around 1.5 million" (Doc. 555-1). How many of either were bought is, in both exhibits, a black bar.
The negotiation emails read like any procurement thread — which is precisely what makes them chilling. In August 2024, a sales manager at Wonder Book, a large used-book dealer, proposed: "We could potentially start with creating a smaller list vs our whole inventory? Perhaps the less common books." Turvey's reply: "sure the less common books are a great place to start… if you have that metadata (ISBN, Title, Author, Publisher, Publication Date, Category, List Price, Our Cost)." His stated reason, in the same thread, was deduplication efficiency — don't sell us what we already bought elsewhere. We flag the register precisely: the vendor proposed starting with uncommon books; Turvey agreed enthusiastically; a strategy of hunting rarity is an inference the documents do not state. What the documents do state is that "less common" was on the table and nobody at Anthropic objected to feeding it into a shredding pipeline.
What happened to them
"There's a machine — there's machines that are built for this kind of work in Germany, I believe, that will cut binds — bindings from books — so cut the glue so that you can then open the book. So it's sheet fed into a scanning machine page by page. A high-resolution image is taken of the pages. There's OCR that is then extracted after the book is, like, sort of — that book is disposed."
— Tom Turvey, deposition at 193

The logistics, from the same pages: books reserved, "taken off shelf if they're on shelf," metadata identified, loaded "in boxes or in Gaylords, put on pallets, put in trucks." Overseas purchases from World of Books arrive "in Illinois, where Datamation is" — the scanning contractor, named under oath. Per shipment, "metadata for the book is handshaked between the scanning vendor and the manifest on the book shipment so that it's one to one. We're sent a copy of what's received." The resulting PDFs go to a site or an Amazon Web Services S3 storage bucket — Turvey wasn't sure which — and from there into what Anthropic calls its "generalized data area."
Access, sworn: "a half dozen, maybe" people at Anthropic plus Datamation staff. Restrictions on what can be done with the scanned books once inside: "I'm not aware of any, no." And the plan for the copies, from a produced document the court quoted: "store everything forever; we might separate out books into categories[, but t]here [wa]s no compelling reason to delete a book."
One more thread from the file deserves daylight. In October 2024, arranging a large purchase through a contact at Amazon — yes, Amazon appears in Anthropic's book supply chain — Turvey wrote: "Can you please let the seller know we're happy to enter into an NDA to preserve and protect their data. I'd like to avoid mentioning us by name but okay to mention we are an AI company you work with building a research library." Under oath, asked whether he was instructing intermediaries to describe the purchases as being for a "research library," he answered: "That's what it says, yes" — and agreed the term was "more of a — sort of a figure of speech," that buying books to train LLMs and building a research library were, as the questioner put it, "one and the same." The concealment wasn't only internal. It travelled down the supply chain.
Distributors shipped pallets; a manifest was handshaked one-to-one against metadata; German-built machines cut the glue; scanners ran; the paper was disposed of; the PDFs went into a private storage bucket reachable by "a half dozen, maybe" Anthropic people plus contractor staff, with no usage restrictions its owner-of-record was aware of — and a written plan to keep everything forever. Every clause in that sentence is sworn testimony or a produced document.
The Licensing Road, Opened and Dropped
Could they have paid authors instead? The record answers with a timeline: Anthropic opened conversations with three of the world's five biggest trade publishers between March and June 2024, then dropped them — not because publishers said no, but because, in Turvey's sworn words, Anthropic "shifted our focus at Project Panama." The wither was a decision with a date.
| Publisher | Contact, per sworn testimony | What happened |
|---|---|---|
| Macmillan | Turvey emailed Stefan von Holtzbrinck (whose family group owns Macmillan) in March 2024; a videoconference followed with staff from a subsidiary | Per Turvey's recollection, "something like" — "We're not sure if we are doing any of this right now… we'll look into it for you." Journal talks later produced a signed NDA on a separate matter; on trade books, Turvey never went back: "we had already switched our focus for books to Panama." |
| Penguin Random House | A senior contact, spring 2024 | She signalled a licensing catalogue would be "complicated" and slow. No follow-up from Anthropic: "By the time I would have maybe turned around and thought about reaching out again, we'd already kind of shifted our focus." |
| HarperCollins | Chantal Restivo-Alessi, chief digital officer, June 2024, introduced via a News Corp contact | Open to the idea, considering an author opt-in model, no plan yet. HarperCollins later did license — Turvey's declaration cites the reported deal: certain backlist non-fiction at $5,000 per book, three-year term, author opt-in (reported November 2024; the counterparty was reported as Microsoft). |
And then the question the whole story turns on was put to him directly — "It was easier to just buy the books from a distributor or reseller, correct?" — and answered under oath:
"It was certainly more cost effective, apparently. Certainly easier from a volume standpoint."
— Tom Turvey, deposition at 256
That is the thesis of this article, in the sworn words of the man who ran the program. Asked why the publisher road stayed closed even after a real deal appeared on it, he was equally plain: "now that we've pivoted our — our plans to Panama, there's not a ton of interest in doing deals with book publishers for books." The reason: "Volume, primarily… just in the U.S. alone, there are over 300 members of the Association of American publishers… So that's at least 300 individual deals… I don't have the resources or the time to do that." And if HarperCollins offered Anthropic the same terms tomorrow? "It is highly unlikely… $5,000 per book, which is way beyond what we're willing to pay." Two of the Big Five he never approached at all: he testified he made no business outreach to Simon & Schuster, and Hachette does not appear in his account of the outreach either.
Asked directly why the publisher route ended, Turvey testified: "by the end of, let's say, the summer of 2024, it became clearer that 'going down the licensing books one publisher at a time' route was not going to succeed because of the volume — simply wasn't there." And: "Book publishers broadly became less interesting at that point, I would say, for direct licensing deals individually." His declaration makes the argument at full strength: "there is not a viable market to license book datasets at any scale remotely approaching Anthropic's complete needs," and of the $5,000-per-book deal that did happen — "I would be shocked if any LLM company had licensed a significant amount of books at prices like that."
Hold the two documented facts next to each other, because this is the crux and it needs no embellishment. Fact one: a real licensing market existed — a Big Five publisher signed a per-book AI licensing deal, with author consent built in, at a knowable price. Fact two: Anthropic's program went the other way, to used-book pallets at a few dollars a volume, where the author's share of a used sale is zero. The court noticed the same pair, citing plaintiffs' expert evidence that "another major technology company soon did" reach a licensing agreement "with one major publisher." Whether $5,000 a book across millions of books was economically survivable is a genuine argument — Turvey makes it under oath, and we've printed it. What is not arguable is the sequence: the priced route existed, and the shredder was cheaper.
At the one reported market price — $5,000 per licensed book — five million books would cost $25 billion. At used-book prices, the court's phrase is "many millions of dollars" for millions of volumes: dollars per book. That gap, three orders of magnitude wide, is the entire documented economics of this story. The sequence, not the motive: the CEO called lawful buying a "slog" in January 2023; the market price appeared in November 2024, twenty-two months later; and by then, per Turvey's testimony, the strategy had already pivoted to Panama.
What Anthropic said afterward
Both companies have issued exactly one sentence about this. Anthropic's, from deputy general counsel Aparna Sridhar after the settlement and the Panama disclosures, as reported by Euronews:
"The issue we settled on was about how some materials were acquired, not whether we could use them to develop [AI]."
Tested against the record, as this piece promised to do with every corporate statement: it is accurate about what the $1.5 billion settled — acquisition, not use — and it is silent on everything else this file contains. It does not mention the millions of purchased books destroyed, the codename, the instruction that the work "should not be shared with anyone outside Anthropic," the licensing conversations opened and dropped, or the catalogue of what was cut apart. A sentence can be true and still be a wall. Note also its grammar of agency: things were "acquired," an issue "was settled." Nobody, in that sentence, does anything.
Follow the Money: The Forensic Ledger
The dollars in this story fall into three kinds, and the reader is owed the boundary between them: figures stated in a primary document, figures derived by arithmetic we show, and figures that exist but are hidden — redacted in an exhibit, sealed on the docket, or never disclosed at all. Everything below is sorted into those three bins. Nothing is estimated.
Stated in the record
| Item | Amount | Source |
|---|---|---|
| Settlement fund, non-reversionary, plus interest | $1,500,000,000 | Doc. 680 at 4, 22 |
| Works on the Works List | 482,460 | Doc. 680 at 4 |
| Works claimed by April 16, 2026 | 440,490 (91.3%) | Doc. 680 at 4 |
| Valid opt-outs | 350 valid opt-outs, spanning 1,802 works | Doc. 680 at 6 |
| Objections or comments filed | 54 | Doc. 680 at 6 |
| Attorneys' fee first requested (20% of fund) | $300,000,000 | Doc. 680 at 14, note 7 |
| Attorneys' fee finally requested (12.5%) | $187,500,000 | Doc. 680 at 14 |
| Attorneys' fee awarded | $101,561,111 — "nearly 6.8%" of the fund; 3.75 × a lodestar of $27,082,963; 10% withheld pending a post-distribution accounting | Doc. 680 at 15–18, 22 |
| Attorney hours credited, actual plus projected to February 2027 | 34,381.6, at hourly rates of $550 to $2,250 | Doc. 680 at 16 |
| Litigation expenses reimbursed | $2,635,197.46 | Doc. 680 at 19, 22 |
| Cost reserve for administration | $18,220,000 (administrator's own estimate about $15 million) | Doc. 680 at 3, 19, 22 |
| Service awards | $15,000 × 3 = $45,000 ($50,000 each was requested) | Doc. 680 at 21–22 |
| Court's per-work figure | "about $3,000, less costs and fees" — "four times" the $750 statutory minimum | Doc. 680 at 4, 17 |
| Reported licence price for a book, HarperCollins backlist non-fiction, three-year term | $5,000, split 50-50 with the author | Turvey decl. ¶36; Authors Guild, Nov. 19, 2024 |
| Anthropic's annual revenue, per the record before the court in June 2025 | "over one billion dollars" | Doc. 231 at 1 |
| Anthropic's annualized revenue run rate, end of July 2026 | $65 billion (a run rate, not audited revenue) | TechCrunch, Aug. 17, 2026, citing Bloomberg |
| Print acquisition spend | "many millions of dollars" / "tens of millions of dollars" | Doc. 231 at 4 / Washington Post, Jan. 27, 2026, paragraph 2 |
Derived — our arithmetic, formula shown
| Quantity | Result | Formula |
|---|---|---|
| Gross fund per listed work | $3,109 | $1.5 billion ÷ 482,460 |
| All court-approved deductions | $122,461,308 — 8.2% of the fund | fee + expenses + reserve + service awards |
| Net distributable, before interest | $1,377,538,692 | $1.5 billion − deductions (assumes the reserve is fully spent) |
| Net per claimed work | ≈ $3,127 | net ÷ 440,490 — the plan pays valid claimants pro rata, so the unclaimed 8.7% flows back to those who claimed; assumes every claim is valid |
| Author's share under the default 50/50 split | ≈ $1,564 per work | net per claimed work ÷ 2 — the default split applies to "non-education works only" (Doc. 680 at 5; Authors Guild) |
| Effective fee award per credited hour (the award includes a 3.75× multiplier) | ≈ $2,954 | $101,561,111 ÷ 34,381.6 |
| Fee cut from the first request | $198,438,889 | $300,000,000 − $101,561,111 |
| Licence price versus settlement allocation, per book — not like for like: a three-year licence against a one-time settlement share | 1.6× | $5,000 ÷ $3,109 |
| Licensing counterfactual — Books3 tranche | $983 million | 196,640 × $5,000 |
| Licensing counterfactual — the class works | $2.41 billion | 482,460 × $5,000 |
| Licensing counterfactual — five million books | $25 billion | 5,000,000 × $5,000 |
| Settlement as a share of the July 2026 run rate | 2.3% — about 8.4 days of revenue | $1.5 billion ÷ $65 billion; × 365 |
| Settlement against the annual revenue in the June 2025 record | at most 1.5× the order's $1 billion lower bound — exact ratio not determinable | $1.5 billion ÷ "over one billion"; the order gives a floor, not a figure |
| Implied cost of a print book | single-digit to low-double-digit dollars — bounds, not a price | "tens of millions" ÷ "millions"; two reported magnitudes, no unit price in the record |
Put the per-book figures side by side — different instruments, a three-year licence against a one-time settlement share against a used-copy purchase, so an order-of-magnitude reading only — and the economics of this story fit on one line. A used copy cost Anthropic dollars, of which the author's share is zero — a used sale pays no royalty. A licensed copy, at the one reported market price, cost $5,000, of which $2,500 reached the author by default. A pirated copy, priced by the court, is settling at about $3,100 gross per listed work — about $3,130 net per claimed work, of which roughly $1,560 reaches a trade author under the default split. The statutory floor for ordinary infringement is $750. And the number Anthropic itself was willing to pay per book is the one figure on that line the public cannot read: Turvey testified that "there's some unit cost there that is an upper bound that we might be willing to pay," produced by "a rather complicated formula that I am not fully in the weeds on"; when the examiner summarized it as attributing "some amount of money you are willing to spend to get the next [figure redacted] tokens," he did not dispute the framing — and every number in that passage is redacted.
The record also settles a question of whether any money ever reached a publisher for this program. Under oath: "Q. Well, you haven't paid any publisher directly for any books, correct? A. Is that true? I believe — yeah, that's probably true. Yeah." (Depo. at 258.)
The fee fight, and what it means for timing
Class counsel first sought 20% of the fund, $300 million, then 12.5%, $187.5 million. Judge Araceli Martínez-Olguín ruled that a percentage award on a fund this size "would result in a windfall," priced the work by hours instead, and awarded $101,561,111 — "nearly 6.8%" — with 10% held back until the parties account for how the money was actually distributed. The order also cut the named plaintiffs' service awards from the $50,000 each requested to $15,000. On August 18 and 19, 2026, three law firms filed two notices of appeal to the Ninth Circuit: the publishers' coordination counsel appealing from the judgment as a whole with a stated focus on the fee order, and a second firm appealing "solely relating to" the fee order as it pertains to that firm — a notice that states it "does not seek to alter the total amount of attorneys' fees awarded." Whether those appeals delay the payment date is a question for the settlement administrator, and we have not found it answered in the record — so we do not answer it either.
Anthropic's per-book ceiling (redacted, depo. 251–253). The Baker & Taylor agreement volume (redacted, depo. 191). The scanning line's daily throughput (redacted, depo. 192). The totals in the go/panama-costs tracker (linked from the Panama document, never produced publicly). The exact print spend (the purchase-program section of Turvey's declaration, ¶¶19–27, is redacted). Anthropic's own legal bill. Interest earned on the fund. And from Amazon: every single number. Each of those figures exists on a spreadsheet somewhere. None of them is in the public record, and this piece prints none of them.
Amazon's Warehouse and the Accountability Gap
Everything known about Amazon's program fits in one paragraph, and that is the finding. There is no court file, no named executive, no stated purpose, no disclosed count — only a tracked shipment, worker accounts, and a statement that does not answer any question a reader would ask. For a company that began as an online bookstore, the silence is itself a document.

What the reporting established: 404 Media's tracker ended at VGT3 — North Las Vegas, per the Review-Journal — a facility whose team emblem is a dinosaur holding a book. A worker told 404 Media (single source, and we label it so) that the site's whole job is receiving bulk book shipments; that a machine "just comes down and slices" the bindings; that roughly 20–25 scanners run; that finished pages are dumped loose into cardboard shuttles; and that intake has included library liquidations, government documents, multilingual pallets and rare volumes. Workers were told the operation was "for Kindles" — the worker doubted it. Orders trace back to 2024 (Inc.).
What Amazon has said, in full, as an exhibit: "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use." No mention of AI. No mention of scanning. No mention of what happens to the books. Snopes reports that Amazon had previously denied engaging in destructive scanning; we could not retrieve the original denial's exact wording (two hosts refused our requests — logged in the could-not-look table below), so we carry that fact at Snopes's strength, not our own.
Two more facts complete the picture. First, from the Anthropic file: when Anthropic wanted books through an Amazon channel in October 2024, an Amazon representative facilitated the conversation — the two stories touch. Second, the org chart: Amazon's frontier-AI group was led by SVP Rohit Prasad until end-2025, then by Peter DeSantis (CNBC) under CEO Andy Jassy. We state that as structure, not as sourced involvement — because no document ties any of them to VGT3. That absence is not our failure to look. It is the difference between a company that has been through discovery and one that has not yet been.
Amazon's documented history with destroying and revoking its own inventory did not begin in 2024. In 2009 it remotely deleted lawfully-questioned copies of Orwell's 1984 from customers' Kindles mid-read (NPR); its CEO apologized, and a student's lawsuit settled. In 2021, ITV filmed a UK warehouse marking 124,000+ items "destroy" in a single week — "books galore" among them (ITV News investigation, June 21, 2021, as reported by the Irish Times). Same company, two decades, three documented destruction programs. Pattern stated; motive not asserted.
Both companies' records extend past books, and the fair-comment rule cuts both ways — so here is each side of the ledger. In May 2025, Dario Amodei told Axios that AI could eliminate half of entry-level white-collar jobs within one to five years and push unemployment to 10–20%: "Cancer is cured, the economy grows at 10% a year, the budget is balanced, and 20% of people don't have jobs." Sequence, not motive: the executive whose acquisition strategy is documented above also publicly forecasts the labour damage of the product it fed. In his favour, and printed here because a one-sided ledger is a broken one: in June 2025 he publicly opposed a proposed ten-year freeze on state AI regulation as "too blunt," arguing for a federal transparency standard instead. Amazon's parallel record is in the box above — the 2009 Kindle deletion, the 2021 Dunfermline destruction filmed by ITV.
It Is Not Two Companies. It Is a Market.
Anthropic is documented because a court forced it and Amazon because a tracker found it — but the booksellers who supply this trade describe an industry-wide buying wave with brokers, NDAs and anonymous buyers. The two named programs are the part that has been dragged into the light, not the extent of the practice.
Scott Brown, a rare-book dealer in Portland, Oregon who writes the trade newsletter Dispatches from the Rare Book Trade, told the Anchorage Daily News the discussion has fixated on Anthropic only because its documents became public: "It's way more than Anthropic." His estimate — and we label it as his estimate, not a measurement — is that tens of millions of books are being destroyed industry-wide, with the orders now reaching independent bookstores. His specific alarm is about titles "where there might only be five or 10 copies available for sale in the entire world," signed editions among them, and about what a scan does not preserve: "AI doesn't really remember the book. There isn't a database of it."
That last point deserves a precise correction, in Brown's favour and against the companies. It is true of the models — weights are not a library. It is not true of the corporate archives: Anthropic's own plan, in the court record, is to "store everything forever." So the accurate statement is worse than his. The book often still exists — as a private PDF in a corporate vault, in a bucket half a dozen people can reach. What is gone is the public copy.
Two more mechanisms, both documented, show the market's shape:
- Brokers with NDAs. ISBNdb, a bibliographic-data company, now brokers high-volume book acquisitions for AI developers — orders reported up to a million volumes, with non-disclosure agreements shielding the buyer's identity. Concealment has become a service you can purchase.
- Anonymous bulk requests, worldwide. Pieter de Vries, an antiquarian bookseller in Haarlem, received an emailed request for 3,001 titles — mostly 2020–21 imprints from Elsevier, Wiley, Routledge, Oxford University Press and Emerald, to be shipped to China — from an entity called 2077AI. He assumed it was "spam or phishing" and ignored it; Fortune reported that several other Dutch antiquarians received the same request and did the same. 2077AI did not respond to Fortune's requests for comment.
The Dutch booksellers' reaction is the most human fact in this investigation. Faced with a genuine industrial-scale request to buy their inventory, the trade's instinct was that no real buyer could possibly want this — so it must be a scam. And an anonymous bookseller quoted by Novara Media, who took the money and still didn't like it, put the trade's position in one line: "I don't like the end-use, and I don't like that uncommon books are being pulped."

Two companies are documented. A broker sells anonymity by the million-volume order. Booksellers on two continents report the buying wave reaching independent shops. And no regulator anywhere counts books bought, scanned or destroyed for AI training. The measured floor is what this article could verify; the ceiling is unknown, and unknown by design.
Is Any of This Legal?
Mostly yes — and the ruling that made it so came after the shredders were already running. In June 2025, Judge Alsup split the case three ways: training on books is fair use; buying a print book, scanning it and destroying the original is fair use; downloading pirated copies never was. The $1.5 billion settlement priced the third. The first two are one district judge's rulings on one company's facts, not binding precedent — and they are the rulings every company in this trade now reads first.
On the destruction step, the order's logic is narrow and worth quoting: "the mere format change was a fair use," because Anthropic already owned each copy — the law "entitle[d] Anthropic to 'dispose[ ]' each copy as it saw fit" under the first-sale doctrine — and the scan "added no new copies": one digital file replaced one destroyed book, kept internally. The court leaned on the microfilm-era Texaco dictum and the Google Books precedent. On the piracy, the language is unsparing: "Pirating copies to build a research library without paying for it" was infringement; there is no fair-use right "to steal a work you could otherwise buy."
Our legal reading, labelled as ours: because the ruling turned on the digital copy replacing rather than adding to the print copy, destruction became the legally safest configuration — keep the book and the scan, and you are arguably holding two copies; burn the book, and you are holding one. Nothing in the order requires destruction. But after June 23, 2025, a company lawyer designing a book-scanning pipeline has an additional legal reason to specify a one-copy replacement structure — which, in practice, is the shredder. That is the state of the law your feed never explained.
And the chronology matters more than any editorial: the Panama document was updated April 13, 2024. Amazon's orders trace to 2024. The ruling came June 23, 2025. The destruction was not caused by the ruling — it was ratified by it, on grounds no produced document states, more than a year after the choice was made.

For Canadian readers asking the question our harvest found them asking — "is it illegal to burn books in Canada?" — the answer is no, not if you own them. Canadian law, like American law, regulates copying and distribution, not what an owner does with a lawfully purchased physical object. The fights above are copyright fights. There is no statute of the shelf.
In Bartz, one district court, June 2025: training on the books at issue — fair use. Buying, scanning once and discarding the print copy — fair use. Piracy — $1.5 billion. Destroying books you own: legal in Canada and the US, always has been. The law's one hard line in this whole story protected the copyright, not the book.
Was Your Book Taken? How to Check
Two public instruments exist, and both take under a minute. The official one is the settlement administrator's lookup at anthropiccopyrightsettlement.com, which searches the certified class list — 482,460 works, each with an ISBN or ASIN and a timely US copyright registration — by title, author, publisher or ISBN. The unofficial one is the public Books3 filename roster — 196,611 entries, mostly "Author - Title" — which covers the first pirated tranche.
Know what the official list is and is not. It is the compensable subset: pirated works that met the settlement's registration criteria — winnowed from roughly seven million downloaded files. It is not a list of the destroyed print books, and it is not Anthropic's full ingestion catalogue. Note also how it is published: the complete works list was filed with the court under seal (the September 24, 2025 notice attaches it as a sealed exhibit), and the public interface is a captcha-gated search box. You may query the list one title at a time; you may not have the list.
The service facts, dated: the claims deadline was March 23, 2026; Judge Araceli Martínez-Olguín granted final approval and entered judgment on July 20, 2026 (Doc. 680); by April 16, 2026, 440,490 of the 482,460 works — 91.3% — had been claimed; the court's own per-work figure is "about $3,000, less costs and fees," and our arithmetic on the order's deductions puts the net at roughly $3,127 per claimed work before interest (Authors Guild; Copyright Alliance). Payment-status questions go to the administrator through the settlement site. One more term worth knowing: the settlement requires Anthropic to destroy the pirated files themselves, with written certification. The paper was destroyed to make files; the stolen files are destroyed by court order. In this story, destruction is the one constant.
The Lists Exist — and No One Will Publish Them
The question readers ask most — which books, exactly? — has a precise, documented, maddening answer: complete title-level records exist inside the companies, and none has ever been published. This is not speculation about hidden knowledge. The record describes the ledgers being built.
- Anthropic's catalogue — court-documented, and promised to the plaintiffs. The order: Anthropic "created its own catalog of bibliographic metadata for the books it was acquiring." The deposition: per-shipment metadata "handshaked… so that it's one to one," with Anthropic "sent a copy of what's received." And the sworn interrogatory answer of December 23, 2024: Anthropic "will produce a document showing the title, author, ISBN, and year of publication of all books from the referenced 'book-scanning' project used in training commercially available and generally accessible Claude models" — adding that "its normal practice is to scan each listed book a single time" (Doc. 554-30, Interrogatories 11–12). So the list is not merely inferred from a database in San Francisco. Anthropic undertook, under Rule 33(d), to hand it to the plaintiffs' lawyers under a protective order — and the public has never seen it.
- The class list — sealed. The 482,460-work list is a court exhibit, filed under seal, publicly reachable only through a search box.
- Amazon's intake — inferred, labelled. A worker describes barcode-scanning every arriving book. If those scans are retained — an inference from a single account, and we say so — a ledger exists in Seattle too.
- The reconstruction — ours, public. What no company will publish, the catalogues partially replace: the Books3 roster (196,611 filenames) and the February 2021 LibGen catalogue (5,221,129 records) are public artifacts. Between them, they name most of the pirated universe. The destroyed-print list has no public substitute at all.
So the demand this piece exists to make is small, specific and entirely within two companies' power by Friday: publish the catalogues. Not the scans — the metadata. Titles, authors, ISBNs, counts. Anthropic's own court filings prove the file exists and prove that releasing bibliographic metadata harms no one — the company itself downloaded LibGen's catalogue as "a separate catalog of bibliographic metadata," because facts about books are just facts. A civilization is owed, at absolute minimum, the index of what was fed through its shredders.
The Canadian Shelf
How many Canadian books were in the pile? Nobody had ever counted, so we did: at least 9,246 records in the pirated LibGen catalogue carry unambiguous Canadian publishing signals — a University of Toronto Press or Presses de l'Université de Montréal imprint, a Dundurn or House of Anansi line, or an unambiguous Canadian city of publication. That is a hard floor, and the honest method is the story of why.
The catalogue is a bad witness for nationality: 83% of its non-fiction rows record no city of publication at all, and its fiction table has no city column, period. Publisher names are populated in only 51–69% of rows. And city names deceive — "London" in a book catalogue is nearly always England, "Victoria" is usually Australia — so we excluded every ambiguous city by design, London, Ontario included. We counted only what cannot be argued with: 7,005 non-fiction and 2,241 fiction records, 9,246 in all, method and script published, samples eyeballed row by row. The true Canadian number is some multiple of the floor. The floor alone is nine thousand books.
Open the sample file and Canadian intellectual history looks back at you: H. S. M. Coxeter's The Fifty-Nine Icosahedra (University of Toronto Press, 1951) — Coxeter, the century's great geometer, a Toronto professor for six decades. The Canadian Mathematical Society's Olympiad volumes, published in Ottawa. The Université de Montréal's mathematics series. École Polytechnique engineering texts. Dundurn's popular history. These aren't abstractions in a seven-million-row ledger; they are the national shelf, and they were in the pirate library Anthropic downloaded in June 2021, per the court's own account of what was taken.

Two Canadian footnotes, carried at exact strength. Press reports have placed Zoom Books, a British Columbia bulk bookseller, in the AI supply chain; Turvey's sworn list of Panama's suppliers does not include it, so we report the discrepancy rather than the claim. And no Canadian regulator — no regulator anywhere, per Snopes — counts books bought, scanned or destroyed for AI training. The 9,246 figure is a count of catalogue records in the pirated LibGen tranche; it says nothing about the print pile. If Canadian books are being pulled from the used market and cut apart in Illinois or Nevada, the count of them exists only in the buyers' private databases — no public floor for that number exists, including here. For what Canadians can do about supply chains that treat this country as an extraction zone, our Buy Canadian investigation maps the levers that actually move; for the sibling story of how Canadian dollars flow into American content pipelines, see the U.S. media bill.
9,246 records in the pirated LibGen catalogue carry unambiguous Canadian publishing signals — counted by us, reproducible by you, and an undercount by design. That is the Canadian floor for the pirated tranche. The Canadian share of the print books Anthropic bought and cut apart is a different number, and it is unknown: it exists only in the company's catalogue.
Has This Happened Before? 2,200 Years of Destroyed Books
Every previous mass destruction of books in the historical record falls into one of three shapes: destroy what the book says, destroy the loser's property, or destroy the carrier for administrative convenience. Knowing which shape 2024–26 belongs to — and where it breaks the pattern — is the difference between an analogy and an understanding. So before the shapes, the ledger: each event with its count, its mechanism, who supplied the stated motive, and what happened to the text.
| Event | Date | Volumes (sourced) | Mechanism | Motive on record — whose words | Fate of the text |
|---|---|---|---|---|---|
| Qin proscription of histories | 213 BCE | No count survives | Fire, by decree | State decree, as reported more than a century later by Sima Qian — details doubted by modern historians (record) | Annihilated by intent |
| Sack of Baghdad; the House of Wisdom | 1258 | No count survives; the "Tigris ran black with ink" line is a report from a 16th-century chronicle, not a count | Conquest | None stated — the books were the loser's property | Lost by indifference |
| Maní auto-da-fé, Yucatán | 1562 | 27 Maya books (EBSCO) | Fire | The destroyer's stated reason, as EBSCO quotes it: "superstitions and falsehoods of the devil" | Annihilated by intent |
| German student burnings | May 10, 1933 | About 20,000 volumes in Berlin alone; burnings in more than 20 university towns and cities (USHMM) | Public bonfires | The destroyers' own proclamations | Annihilated by intent — the burning was the message |
| Jaffna Public Library | Night of May 31–June 1, 1981 | 95,000–97,000 volumes — the lower figure in contemporary accounts, the higher from the library's retired chief librarian (Tamil Guardian) | Arson by state security forces and state-sponsored mobs (Tamil Guardian) | None from the actors; ethnic-cultural erasure is the historians' characterization | Annihilated — palm-leaf manuscripts with no other copy |
| National and University Library, Sarajevo | Night of Aug. 25–26, 1992 | "About 2 million volumes" (UNESCO) | Deliberate shelling; "fire ignited by grenades" (UNESCO) | None from the gunners; erasure of a multi-ethnic record is scholarship's assessment | Annihilated |
| The microfilm era, US and UK libraries | "For fifty years" per the publisher; boom of the 1980s–1990s (record); British Library sales 1999 | "Hundreds of thousands of historic newspapers" (Baker, Double Fold) plus thousands of books | Guillotine, camera, discard | The institutions' own published rationale — space and preservation; Baker argues it was largely pretext | Text kept, publicly, in a worse format |
| Stripped mass-market paperbacks | Ongoing, by contract | No public count | Cover stripped for credit, body pulped | Commercial — written into the returns contract | Text unaffected — other copies exist |
| Google Books | 2004– | "Over 40 million books" (Turvey decl. ¶7) | Non-destructive camera rigs; volumes returned to libraries | Stated publicly — search and access | Book kept; text indexed, snippets shown |
| Internet Archive | Ongoing | Millions of items; "over 500,000 books" removed from lending after the 2024 ruling (IA) | Non-destructive scanning; physical books kept in shipping containers in Richmond, California (IA) | Stated publicly — preservation and lending | Book kept; text lent one copy at a time |
| Anthropic, Project Panama | 2024– | "Millions" — the exact count is in Anthropic's catalogue and is not public | Spine cut, sheet-fed scan, paper disposed | None stated for the destruction step in any produced document; acquisition motive documented (training data, cost) | Text kept privately, "forever"; the public copy gone |
| Amazon, VGT3 | 2024– | Undisclosed | Spine cut, scanned, pages discarded, per workers | None stated; the one corporate sentence names no purpose | Unknown — no document |
Two things the table shows that no paragraph could. First, on volume: no cross-era comparison is printed here because the AI rows have no number — and the reason they have no number is not that nobody counted. Anthropic's sworn interrogatory answers of December 23, 2024 state that it scans "each listed book a single time" and undertake to produce the title list of what it scanned. Among the cases compared here, Anthropic's is the only one where the party that destroyed the books is on the sworn record as holding a title-level list of what it destroyed, and has not published it. Landa's twenty-seven is the number history was given; Anthropic's is on a server. Second, on motive: the censors of 1562 and 1933 and the microfilm librarians put their reasons on the record; the conquerors and arsonists never did. The AI rows are the only ones where a destroyer who was neither a conqueror nor an arsonist — a buyer, in peacetime, with a cost tracker — gave no reason for the destruction step.
Shape one: erasure — destroy what the book says
Qin Shi Huang's chancellor ordered proscribed histories burned in 213 BCE (the account first appears more than a century later, in Sima Qian, and modern historians doubt its details — we carry the caveat). At Maní in 1562, Friar Diego de Landa burned 27 Maya books — the one entry on this list where the destroyer's stated reason survives in his own words, as EBSCO quotes them: "superstitions and falsehoods of the devil." Centuries of a civilization's writing, gone in an afternoon. On May 10, 1933, German students burned "un-German" books in more than 20 university towns and cities — about 20,000 volumes in Berlin alone, some 40,000 people watching. On the night of May 31, 1981, Sri Lankan state security forces and state-sponsored mobs burned the Jaffna Public Library — 95,000 to 97,000 volumes and irreplaceable Tamil palm-leaf manuscripts; historians characterize it as ethnic-cultural erasure. In August 1992, gunners in the hills deliberately shelled Sarajevo's National Library; about two million volumes burned, an act scholarship assesses as an attempt to destroy the documentary record of a multi-ethnic Bosnia. The common thread: these destructions were public on purpose. The burning was the message.
Shape two: conquest — the books are just the loser's property
Baghdad, 1258: Hulagu's army destroyed the House of Wisdom and the city's dense fabric of libraries; the Tigris was said to have run black with ink — a line from a 16th-century chronicle, not a count (record). The books died of indifference, not doctrine.

Shape three: administration — destroy the carrier, keep something
This is the lineage almost nobody knows, and it is the one that fits. For fifty years, in the publisher's account of Baker's book, and at a peak in the 1980s and 1990s, American libraries — the Library of Congress included — guillotined hundreds of thousands of bound newspaper volumes, and thousands of books, so the pages would lie flat under the cameras, then discarded the originals; in 1999 the British Library auctioned off runs of its American newspaper collection. Nicholson Baker documented it in Double Fold (2001) and cashed in his retirement account to save one twenty-ton archive (record). Same machine as Project Panama — guillotine, camera, discard. The libraries' stated justification was space and cost, and Baker's own argument was that the access rationale was largely pretext; Anthropic's reason for the destruction step is unstated anywhere in the record. The publishing industry itself runs a standing version: unsold mass-market paperbacks are stripped of their covers for credit and pulped by the ton, every year, by contract (record).
Three scanning projects, three rulings
The nearest precedents are not bonfires at all. They are the two other industrial book-scanning projects of the digital era, and American courts have now ruled on all three. Read the outcomes in a row.
| Project | What it did to the book | What it did with the text | Ruling |
|---|---|---|---|
| Google Books (from 2004; Turvey ran partnerships) | Scanned over 40 million volumes without destroying them; library copies returned | Public search index; snippets shown; full text withheld | Fair use. Second Circuit, Oct. 16, 2015; Supreme Court declined review, April 18, 2016. A proposed $125 million settlement had been rejected in 2011 (record). |
| Internet Archive | Scanned books it owned; keeps physical copies in shipping containers at its Physical Archive in Richmond, California | Lent one digital copy per physical copy owned; 127 works at issue | Not fair use. S.D.N.Y., March 24, 2023; affirmed by the Second Circuit, Sept. 4, 2024; no further appeal (record). Over 500,000 books withdrawn from lending. |
| Anthropic, Project Panama | Bought millions of used books, cut the spines, scanned, discarded the paper | Private archive, "store everything forever"; used to train a commercial model; no public access | Fair use for the buy-scan-destroy step and for training. N.D. Cal., June 23, 2025. Piracy of the digital tranches: $1.5 billion settlement, final judgment July 20, 2026. |
Stated as sequence, because the three courts applied different law to different facts and this piece will not pretend otherwise: the project that kept every book and let the public read the text lost; the project that kept every book and showed the public snippets won; the project that destroyed the books and kept the text for itself won. Whatever the doctrinal reasons — and they are real; the Second Circuit found the Archive's lending non-transformative, while Judge Alsup found training "exceedingly transformative" — the sequence is the record, and what any company lawyer takes from it is an inference we label in the section below, not a finding. Asked under oath whether Google Books and Project Panama differ, Turvey answered: "I mean, I suppose the — the intent and benefit of the digitization is slightly different, and both are —" and was cut off by simultaneous colloquy. (Depo. at 148.)
Where 2024–26 fits — and where it doesn't
Mechanically, the AI programs are Shape Three: cutters, cameras, cost logic, no bonfire, no crowd. In scale, they are Shape One: millions of volumes, a national-library order of magnitude. But on the axis that defines all three shapes — what happens to the text — they are something the record has no precedent for. The censors wanted the text annihilated. The conquerors didn't care. The microfilmers kept the text public, in a worse format. Here, the produced documents describe a fourth configuration: the text preserved with a plan to "store everything forever," held privately, reachable by "a half dozen, maybe" Anthropic staff plus the scanning contractor's people, while the public copies went into recycling bins — under an instruction that the fact of the work "should not be shared with anyone outside Anthropic."
One more exhibit bears on what "preserved" means here, and it is carried at exactly the strength of its source. On March 11, 2024, a professor at ETH Zurich wrote to Anthropic's user-safety and disclosure addresses that his group's experiments suggested the Claude 3 Opus model "reconstructs (copyrighted) training data with little prompting effort" and "generates texts by mixing together verbatim pieces of texts from different sources"; the next day the chain had reached a deputy general counsel, who replied to the company's outside public-relations agency: "Thank you! We're aware of this. Will add press@ to the other thread." (Doc. 555-6.) Those are the researchers' reported findings, not adjudicated facts, and they concern a model rather than the scan archive. But they close the loop the rare-book dealer Scott Brown opened when he said "AI doesn't really remember the book." On the record, the book is in a vault — and, per an outside laboratory's report that the company's legal department said it was aware of, pieces of it may be in the product.
Our reading — flagged as ours, because no document states a posture toward books and we will not invent one: the censors and the microfilmers said why they destroyed. Landa's reason was recorded; the students proclaimed theirs; even the microfilm librarians published their justifications. These companies' documents state no reason for the destruction at all — only cost trackers, token forecasts, dedupe scripts and a secrecy clause. If you want the human-nature throughline, it is the oldest one in economics: when a resource becomes scarce and valuable — and verified pre-2022 human writing is now exactly that — whoever can enclose it, will, unless something stops them. The enclosure of common land looked like hedges and surveyors' stakes, not armies. This one looks like Gaylord pallets and a dinosaur logo. That is an analogy, not a finding. The finding is the sequence: bought, cut, scanned, disposed, retained privately, forever — with the reasons — from the only destroyer on this ledger who bought the books in peacetime and kept a cost tracker — left blank.
We put the pattern of manufactured historical forgetting in a wider frame in 100 Lies That Changed the World, and what happens to a society's trust when records stop being checkable in What Happens When We Stop Believing Each Other?
What the Record Does Not Show
A piece that claims to be built on documents owes you the list of what the documents do not contain. Here it is — the section a press release would never include.
No document states why the paper was destroyed. The acquisition motive is documented (training data; books were "the most cost-effective means to achieve a world-class LLM"). The retention plan is documented ("store everything forever"). The destruction step's rationale is stated nowhere we could find. Three candidate explanations exist, all inference and labelled so: process economics (sheet-feeding loose pages is faster than page-turning intact books — industry mechanics, and a cost tracker existed, but no memo says "destroy because cheap"); legal architecture (one copy replacing another is what the fair-use ruling later blessed — but that ruling post-dates the choice, and no produced legal advice is in the record); and secrecy (you cannot resell or donate millions of de-spined books without the program becoming visible — consistent with the documented concealment instruction, causal link unproven). If a document surfaces stating the reason, we will update this piece and say so.
Could not look — the full table. Found, not-found and could-not-look are different findings; these were the third kind:
| Item | Route tried | Status |
|---|---|---|
| Washington Post's Jan. 27, 2026 investigation (original text) | Direct fetch refused; Internet Archive snapshot located | Paywalled at source; cited via the archived copy and corroborating recaps, flagged where relied on |
| Amazon's earlier denial of destructive scanning, verbatim | Two hosts refused automated requests (403/402) | Carried at Snopes's strength only |
| Four sealed-then-PACER-only exhibits | Free mirrors only; we do not pay for public records | Unread; nothing in this piece relies on them |
| Amazon's book volumes, vendor list, program owner | No disclosure exists to retrieve | Absent by the company's choice |
| Anthropic's print-destruction total beyond "millions" | Court record; WaPo-reviewed filings | The precise number sits in the company's internal catalogue — and in the title list Anthropic undertook to produce to plaintiffs under protective order (Doc. 554-30) |
| Anthropic's per-book valuation ceiling; Baker & Taylor agreement volume; scanning throughput | Turvey deposition, Doc. 554-22, pp. 191–192, 251–253 | Each figure exists in the transcript and is redacted in the public exhibit |
| The go/panama-costs tracker totals | Linked from Doc. 554-21 | Not produced publicly |
| Washington Post text beyond its first two paragraphs | Internet Archive snapshot of Aug. 21, 2026; page data parsed | Paywalled; two paragraphs recovered and saved, the rest carried via Euronews and IBTimes UK recaps, labelled |
| Effect of the August 2026 fee appeals on payment timing | Docs. 682, 683 read in full | Not answered in the record retrieved; not answered here |
And the standing correction ledger: in drafting, our own editor process killed three claims before publication — a "rarity was the shopping criterion" framing (the vendor proposed it, not Anthropic; corrected above), a cross-era volume comparison that likely fails against Sarajevo's two-night toll (deleted; the per-event table above now carries sourced counts instead), and motive language the documents do not support (rewritten as sequence). A second pass on September 6, 2026, made four more corrections before publication: the Jaffna count is printed as 95,000–97,000, both figures as the Tamil Guardian carries them, not a single figure; the microfilm era is dated to the publisher's "fifty years" and the documented 1980s–1990s peak, replacing an unsourced "forty"; the settlement's final approval is dated July 20, 2026 from the order itself, where press reports said July 21; and two figures that had been attributed to the Washington Post directly — the 500,000-to-two-million vendor proposal and the 4,000 pages — are now labelled as reaching us through recaps, because only two paragraphs of the Post's text could be retrieved. We publish that we caught them, because a correction you never see is just a smaller version of a sealed exhibit.
Frequently Asked Questions
Are AI companies really scanning and destroying millions of books?
Yes. Federal court records show Anthropic bought millions of print books, cut them from their bindings, scanned them and discarded the paper; an internal document unsealed in January 2026 calls this Project Panama. Separately, 404 Media tracked rare books to Amazon's VGT3 facility in the Las Vegas area, where workers describe the same cut-scan-discard process.
Why is Amazon destroying books?
Amazon has never given a reason. Its only statement says it "purchases books through commercial channels to help develop and improve the products and services our customers use" — it does not mention AI, scanning, or what happens to the books. Workers told 404 Media they were told the operation was "for Kindles," which one doubted. The absence of any named executive or stated reason is itself a documented finding.
How many books did Anthropic pirate?
Over seven million copies, per Judge Alsup's June 2025 order: 196,640 books from Books3 (downloaded January–February 2021), at least five million from LibGen (June 2021), and at least two million from Pirate Library Mirror (July 2022). Our independent count of the February 2021 LibGen catalogue found 5,221,129 records, consistent with the court's figure.
Did Anthropic use my book?
Check the official settlement lookup at anthropiccopyrightsettlement.com, which searches the 482,460-work class list by title, author, publisher, ISBN or ASIN. The Books3 filename roster (196,611 entries) is also public on GitHub. The full works list itself was filed under seal; the public interface is the search box.
What is Project Panama?
Anthropic's internal name for its book-destruction pipeline. The unsealed planning document reads: "Project Panama is our effort to destructively scan all the books in the world," explains the codename exists because "we don't want it to be known that we are working on this," and links to internal cost and token-forecast trackers. Its listed owner is Tom Turvey. Last major update: April 13, 2024.
Who is Tom Turvey?
Anthropic's Vice President of Product Partnerships, hired February 2024, and the listed owner of the Project Panama document. By his own sworn declaration he spent 20 years at Google, helped create Google Books in 2004 — a project that scanned over 40 million books non-destructively — and concluded thousands of publisher licensing deals before running the program that cut books apart.
What is Books3?
A pirated dataset of roughly 197,000 books assembled from a private torrent tracker, used to train multiple AI systems. The court found Anthropic co-founder Ben Mann downloaded it in early 2021 knowing it was pirated. The full filename list is public on GitHub (196,611 entries), and authors can search it to see whether their work was included.
What is Amazon VGT3?
An Amazon facility in the Las Vegas area — North Las Vegas per the Las Vegas Review-Journal — identified by 404 Media after a tracking device hidden in a shipment of rare books ended up there. Workers describe its sole function as receiving bulk book shipments, machine-cutting the spines, running pages through roughly 20–25 scanners, and discarding the loose pages. Its team logo is a dinosaur holding a book.
Is it legal to destroy books you own?
Yes. A lawfully purchased book is property, and no Canadian or American statute forbids destroying your own copy. The legal fights are about copying, not destruction: Judge Alsup ruled in June 2025 that buying a print book, scanning it once and destroying the original is fair use, because the digital copy replaces rather than adds to the print copy — while downloading pirated copies cost Anthropic a $1.5 billion settlement.
Is it illegal to burn books in Canada?
Not if you own them. Canada has no law against destroying your own books — open-air burning may breach local fire bylaws, but the destruction itself is legal. What Canadian law regulates is copying and distribution under the Copyright Act, not what an owner does with a lawfully bought physical copy.
What books was Claude trained on?
No complete public list exists. The court record establishes the sources: Books3 (196,640 books), LibGen (5M+), PiLiMi (2M+), and millions of purchased-and-scanned print books catalogued internally. Anthropic's sworn interrogatory answers of December 23, 2024 undertake to produce the title, author, ISBN and year of every scanned book used to train its released Claude models — to the plaintiffs, under a protective order, not to the public. The settlement's 482,460-work list covers pirated works with registered US copyrights — searchable, but a subset.
When will the Anthropic settlement pay authors?
The claims deadline was March 23, 2026. Judge Araceli Martínez-Olguín granted final approval and entered judgment on July 20, 2026: a $1.5 billion fund, attorneys' fees cut to $101.6 million (nearly 6.8%), and a per-work figure the court states as about $3,000 less costs and fees. By April 16, 2026, 440,490 of the 482,460 works (91.3%) had been claimed. Two notices of appeal, both focused on the fee order, were filed August 18–19, 2026; payment timing is administered through anthropiccopyrightsettlement.com, where you can check your claim status.
Is LibGen illegal?
Distributing copyrighted books without permission is unlawful in Canada and the US, and Judge Alsup ruled that downloading from LibGen for AI training was not fair use — "a book that could have been bought at a bookstore" — which is what the $1.5 billion settlement priced. The catalogue metadata (titles, authors, ISBNs) is a different thing: facts about books are not copyrightable, which is how this article could count the catalogue without touching a single book file.
How many Canadian books were in the pirated libraries?
At least 9,246 records in the February 2021 LibGen catalogue carry unambiguous Canadian publishing signals — University of Toronto Press, Presses de l'Université de Montréal, Dundurn, or an unambiguous Canadian city of publication. That is a hard floor, not a total: 83% of non-fiction rows record no city, fiction rows record none at all, and ambiguous cities like London were excluded. Method and script are published and reproducible.
The Bottom Line — and the Demand
Strip the adjectives and this is what the file holds. Between 2021 and 2022, Anthropic downloaded over seven million pirated books; the court found that co-founder Ben Mann knowingly downloaded Books3 and at least five million LibGen copies, and that the company knew the PiLiMi tranche was pirated too. In 2023, buying books lawfully was called a "slog" at the top of the company, and the thread was marked resolved. In 2024, a veteran of the world's gentlest book-scanning project was hired to obtain "all the books in the world," opened the licensing road, watched a real per-book price appear on it, and swung the program to used-book pallets and German cutting machines instead — under a codename, with a written instruction that it not be known. In 2025, a judge blessed the destruction, priced the piracy at $1.5 billion, and the industry kept cutting — Amazon's warehouse included, behind one sentence and no names. In 2026, the documents came unsealed, and the only thing still hidden is the one thing everyone asks for: the list.
So we end where the evidence points. Anthropic: publish your catalogue. Your own filings prove it exists, row by row, ISBN by ISBN, and prove you consider bibliographic metadata to be uncopyrightable facts — you downloaded LibGen's. Amazon: name the program, its owner, and its count. A company founded on books that is cutting them apart at industrial scale owes the public at least the arithmetic. Until either happens, the fullest public accounting of what went into the shredders is a court file, two catalogues, one filename roster — and this page, where at least 9,246 Canadian records in the pirated catalogue are finally counted as what they were — the print pile's Canadian share still uncounted by anyone but the company that cut it.

Books survived Qin's chancellor, Landa's bonfire, the microfilm guillotines and the pulping contracts because someone, every time, kept the index. Keep this one.
Every factual claim above traces to the linked record. If you are named in this piece and believe any fact is wrong, or if you hold a document that answers what the record leaves blank — particularly the destruction rationale or the catalogues — write to milad@zeusebikes.ca. Substantiated corrections will be published in place, dated, and noted — the same standard we applied to our own three pre-publication errors, disclosed above.
Buy Canadian: What Actually Works · The Boycott Stops at the Living Room: U.S. Media's Canada Bill · What Happens When We Stop Believing Each Other? · 100 Lies That Changed the World
Milad Ghobadibeygvand, BScN (Western University, 2014), is the co-founder of Zeus eBikes Canada. This investigation was researched and drafted with Claude, an AI system built by Anthropic — one of its subjects; the conflict, the method, and every source are disclosed above, and the counting scripts are reproducible on request.
Visuals created by Playcut.ai
Sources — primary record first (all links verified live at publication)
- Order on Fair Use, Doc. 231 — Bartz v. Anthropic, N.D. Cal. 3:24-cv-05417, June 23, 2025
- The Project Panama document, Doc. 554-21 (Bates ANT_BARTZ_000005492), unsealed Jan. 21, 2026
- "Anthropic Plan for 2023" comment thread, Doc. 554-27 (Bates -000002604)
- October 2022 chat log, Doc. 554-19 (Bates -394794)
- Vendor correspondence, Doc. 555-1 (Bates -035587–91)
- Turvey deposition excerpts, Doc. 554-22 · Turvey declaration, Doc. 127 · Works-list notice, Doc. 422 · full docket
- LibGen metadata snapshot, Feb. 12, 2021 (catalogue only — contains no books) · Books3 filename roster
- 404 Media: the tracker investigation (Aug. 17, 2026) · 404 Media: inside VGT3 (Aug. 26, 2026) · Las Vegas Review-Journal
- Snopes news analysis (Sep. 1, 2026) · Washington Post (Jan. 27, 2026; paywalled — linked via the Internet Archive) · Fortune · Anchorage Daily News · Euronews · Novara Media · TechCrunch · Inc.
- Settlement works-list lookup · Authors Guild guidance · Copyright Alliance · Authors Alliance analysis
- History: Maya codices (EBSCO) · 1933 burnings (USHMM) · Jaffna 1981 · Sarajevo 1992 (UNESCO) · Baghdad 1258 · the microfilm era · stripped books · the 1984 Kindle deletion (NPR) · Dunfermline 2021 (ITV News investigation, via the Irish Times)
- Order Granting Final Approval, Fees and Judgment, Doc. 680 (July 20, 2026) · Notice of Appeal, Doc. 682 · Notice of Appeal, Doc. 683 · Interrogatory responses, Doc. 554-30 · Océano correspondence, Doc. 555-2 · ETH Zurich disclosure chain, Doc. 555-6 · Authors Alliance on final approval · Authors Guild on the HarperCollins licence · TechCrunch on the $65B run rate · IBTimes UK recap of the Post
- Scanning-project rulings and comparators: Authors Guild v. Google · Hachette v. Internet Archive · Internet Archive on the ruling · Internet Archive's Physical Archive · Double Fold (publisher) · Tamil Guardian on Jaffna
- Context: CNBC on Amazon's AI leadership · Axios: Amodei on white-collar jobs · The Verge: the HarperCollins licensing deal



Share:
Best Electric Bikes for Winter in Canada: 23 Rider-Matched Picks