AI Policy

Anthropic Confirmed Destroying Books to Train Claude. Here’s What “Rare” Means.

Snopes confirmed Anthropic sliced spines off books and destructively scanned them to train Claude. A key question: what counts as "rare"? The answer is murky.

LUMIEN4 min read
Anthropic Confirmed Destroying Books to Train Claude. Here’s What “Rare” Means.

Internal documents unsealed in a copyright lawsuit against Anthropic show the company bought books in bulk, sliced their spines, and ran them through destructive scanning to feed its Claude AI model. When Snopes published that finding, a reader flagged that it conflicted with Anthropic's public statement that no "rare or antiquarian books" were scanned. Snopes pressed the company for a definition. The answer: "rare" means collectible objects, not hard-to-find professional or reference titles, a gap wide enough to cover a lot of books most people would call uncommon.

What happened

Detail Fact
Source Unsealed internal Anthropic documents from a copyright lawsuit
Method Spines sliced off books, then “destructively scanned”
Project lead Tom Turvey
Stated priority “Less common books” (term left undefined in documents)
Anthropic’s public claim Did not scan “rare or antiquarian books” (per statement to The Guardian)
Anthropic’s clarification to Snopes “Rare” means collectible-grade items, not hard-to-find professional or reference books
Snopes verdict “Mostly true”

Snopes reporter Anna Rascouet-Paz combed through hundreds of pages of court documents to confirm the practice. Project lead Tom Turvey had agreed to prioritize “less common books,” but that phrase was never defined in the documents. Snopes initially rated the claim about rare books as mostly supported by the evidence.

After publication, a reader pointed out that Anthropic had told The Guardian it had not scanned “rare or antiquarian books.” That appeared to contradict Snopes’ findings, so the outlet went back to Anthropic for a precise definition of “rare.”

What does Anthropic actually mean by “rare”?

Anthropic’s first response to Snopes was the same boilerplate statement it had given The Guardian. Snopes pushed further, specifically asking the company to define its terms. The eventual answer: the company treats “rare and valuable” as a category equivalent to collectibles. Books that are simply hard to find, like specialist professional texts or niche reference works, do not qualify as “rare” under that definition.

In practical terms, a book can be out of print, sold by only a handful of second-hand dealers, and priced well above a typical paperback, yet still fall outside Anthropic’s self-defined protection. That leaves the company with considerable flexibility over which titles it can process.

Why it matters

This story sits at the center of the copyright debate around large language model (LLM) training, the process of feeding a model vast amounts of text so it learns to generate language. Courts are actively examining whether buying physical copies of books and scanning them constitutes fair use or infringement. The documents give plaintiffs in the lawsuit concrete evidence of the physical process Anthropic used, not just an abstract claim about data ingestion.

The definitional question matters too. By controlling what counts as “rare,” Anthropic can narrow the scope of any damage claim or public criticism. If “rare” only means a signed first edition, then almost any book a library would stock is fair game under the company’s own framing.

For businesses thinking about how AI training data is sourced, this case is a useful reminder that public statements from AI labs often leave key terms undefined. That ambiguity tends to resolve in the lab’s favor until someone presses for specifics. Teams exploring AI integration for their own workflows should factor data provenance questions into any vendor evaluation.

What about spotting AI-generated images?

The same Snopes newsletter covered a separate but related topic: how to tell whether an image was made by an AI. Reporter Jack Izzo noted that the old tells, especially garbled text inside images, are less reliable now as models have become more precise with typography.

Two tips Izzo highlighted as most practical:

  1. Browse with active skepticism. Treat unfamiliar images the same way you would an unsourced statistic.
  2. Examine complex shapes closely. Hands, hair, and ears still trip up AI models. Look for fingers that merge, ears that blur into hair, or hands with the wrong number of digits.

Snopes also noted it now uses Google’s SynthID tool (a watermarking and detection system built into Google’s image generators) in most of its AI image debunks, though it does not rely on SynthID alone. We covered a related development when Google Gemini introduced a toggle for visible AI watermarks, which is part of the same SynthID infrastructure.

Our take

The Anthropic book story is a good example of why the first press statement is rarely the whole story. “We don’t scan rare books” sounds reassuring until you ask what rare means, and the answer turns out to exclude most of what a reasonable person would call uncommon. That is not a mistake. It is a deliberate drafting choice.

For anyone building on top of AI models, it is worth asking providers direct questions about training data sources and getting specific definitions in writing. Vague assurances about data quality or legality are easy to issue and hard to enforce. The legal exposure from copyright litigation in this space is still being worked out in courts, and the outcomes will affect every business that builds on these models.

If you want to understand how AI tools fit into your business without inheriting someone else’s legal ambiguity, talk to the Lumien team about what questions to ask before committing to a platform.

Source: Bing News · Anthropic

Frequently asked questions

Did Anthropic really destroy books to train its AI?

Yes. Internal documents unsealed in a copyright lawsuit show Anthropic bought books in bulk, sliced their spines, and ran them through destructive scanning to train its Claude AI model.

What does Anthropic mean when it says it didn't scan rare books?

Anthropic told Snopes that 'rare' means collectible-grade items, like rare antiquarian editions. Hard-to-find professional or reference books do not qualify as 'rare' under the company's own definition, giving it broad room to scan uncommon titles.

How can I tell if an image is AI-generated?

Check complex shapes like hands, hair, and ears, as AI models still struggle with these. Also look for garbled or oddly perfect text. Google's SynthID tool can help detect AI watermarks, but should not be used as the sole method.

What copyright lawsuit is Anthropic involved in over books?

Anthropic is facing a copyright lawsuit in which internal documents were unsealed. Those documents revealed details about how the company physically acquired and scanned books to build training datasets for its Claude AI model.

More from AI