What fair use means for AI training data — why bought books count and downloaded ones do not
Fair use is the doctrine in United States copyright law that allows limited use of a protected work without permission from the rights holder. It is not a list of permitted activities but a four-factor test courts apply case by case: the purpose and character of the use, the nature of the work, the amount used, and the effect on the market for the original. AI training split along an unexpected line. In 2026 US courts accepted that training a model on lawfully purchased books can qualify as fair use, on the reasoning that training extracts statistical patterns rather than reproducing text and is therefore transformative. But acquiring the same books as pirated copies was held to be infringement, and that holding led to a large settlement. The dividing line is not what the model learned but how the copy was obtained. The music publishers' suit filed against Anthropic on August 28, 2026 targets exactly that line
The three lines
- Definition — an exception allowing use without permission. Not a list, but a four-factor court test
- The line — training on purchased books can be fair use; acquiring pirated copies was infringement
- The point — what decided it was not what the model learned but how the copy was obtained
Key questions
- What is fair use?
- **An exception in United States copyright law that permits using a protected work without permission.** The most common misunderstanding first: fair use is **not a list of situations that are allowed.** It is a **test** a court applies, weighing four factors. ① **Purpose and character of the use** — is it commercial, and is it **transformative**, meaning does it create a new meaning distinct from the original? ② **Nature of the work** — factual records receive thinner protection than creative works. ③ **Amount and substantiality used** — how much, and was it the heart of the work? ④ **Effect on the market** — does the use substitute for the original or its licensing market? In practice **factors ① and ④ are usually decisive.** And a caution: Korean copyright law has its own equivalent provision, but with different requirements and a different body of case law — **US rulings do not transfer directly.**
- Is AI training fair use or not?
- **Training alone may qualify; acquiring the data is judged separately.** US court reasoning in 2026 split the question in two. On **training itself**, the transformative argument was accepted: a model does not store sentences to retrieve later but **extracts statistical patterns** across enormous amounts of text, and the result does not substitute for the original work. But **how the books were obtained** is a different question entirely. Downloading pirated copies is itself a reproduction-right violation, and **whatever came afterwards does not undo it.** Courts treated that portion as infringement, and a large settlement followed. Put simply: **if you bought the book and trained on it, there is an argument to have; if you torrented it, you lose before training even begins.**
- So what should an AI company do?
- **On the evidence so far, the direction is: be able to prove where the data came from.** The rulings turned on acquisition, not method. Three practical routes exist. ① **Purchase and licence** — costly, but the least exposed. This is why large licensing deals in news, music and images have been increasing. ② **Public-domain and permissively licensed data** — legally safest, but limited in volume. ③ **Provenance records** — keeping a durable account of what was obtained, from where, and on what terms. That is the starting point of any defence. But **far more is unsettled than settled.** Only a handful of decisions exist, appeals remain, and whether the same reasoning extends to music, video or images has not been tested. This page does **not** claim the doctrine is settled.
In the AI training-data cases, what decided the outcome was not what the model learned. It was how the book was obtained.
To see why that sentence holds, start with fair use.
1. Fair use is not a list
The most common misunderstanding, first. Fair use is not a list of "these situations are fine." It is a test a court runs.
| Factor | What it examines | Practical weight |
|---|---|---|
| ① Purpose and character | Commercial? Transformative? | High |
| ② Nature of the work | Factual record or creative work | Medium |
| ③ Amount and substantiality | How much, and was it the heart? | Medium |
| ④ Market effect | Does it substitute for the original? | High |
Factors ① and ④ usually decide the case. And the word transformative in ① sits at the centre of the AI debate — does the use create a different purpose and different meaning from the original?
Korean copyright law has its own equivalent provision, but with different requirements and a different body of case law. US rulings do not transfer directly.
2. Where AI training split
US court reasoning in 2026 cut a single dispute into two pieces.
| Stage | Conduct | Holding |
|---|---|---|
| ① Acquisition | Books obtained by purchase | No problem |
| ① Acquisition | Books obtained as pirated copies | Infringement |
| ② Training | Training on the acquired text | Fair use available |
The reasoning behind ② is this. A model does not store sentences and retrieve them later. It extracts statistical patterns across enormous amounts of text, and the resulting product does not stand in for the book — that is, the use is transformative and the market-substitution effect is weak.
But ① is a different question. Downloading a pirated copy is itself a reproduction-right violation. What happened afterwards does not undo it.
Hence the summary: bought books count, downloaded ones do not.
3. That line is now being aimed at
On August 28, 2026, Sony Music Publishing, Warner Chappell Music and other publishers sued Anthropic ("Sony, Warner and other publishers sue Anthropic (August 28, 2026)").
One word recurs in the complaint: torrent.
| How the plaintiffs framed it | What it targets |
|---|---|
| "Millions of pirated books via torrent" | ① Acquisition — the side that already lost |
| "Lyrics and sheet music were inside" | Extends the reasoning to music |
| CEO named personally | Aims at wilfulness |
The plaintiffs are not trying to relitigate whether training is transformative. That argument has already been lost once. They are placing the case on the side of a boundary that has already been drawn in their favour.
4. So what are companies doing
The direction the rulings point in is singular: be able to prove where the data came from.
| Approach | Strength | Limit |
|---|---|---|
| Purchase and licence | Least legal exposure | Expensive |
| Public domain / permissive licences | Low risk | Not enough volume |
| Provenance records | The starting point of any defence | Hard to apply retroactively |
This is the background to the growth of large licensing deals in news, music and images. The arithmetic has begun to favour paying for a licence over paying for litigation.
5. Far more is unsettled than settled
This page does not claim the above is settled doctrine. Here is what is open.
- There are only a handful of decisions. Most are trial-level, and appeals remain.
- A different medium may produce a different result. A holding about books does not automatically govern music, video or images. Music in particular raises the reuse of short phrases in ways text does not.
- Outputs are a separate question. Even where training is lawful, a model producing something closely resembling the original is its own problem.
- Jurisdictions differ. The EU has an explicit text-and-data-mining exception; Korea has yet another framework.
- Company data sources are mostly undisclosed. Very few firms publish what they trained on.
6. What is still unresolved
- The opinion — The "purchased is fair use, pirated is infringement" holding was not checked against the original text.
- The settlement — Its amount, and whether an appeal is pending, were not confirmed.
- Music — No case law exists on whether the same reasoning applies to compositions.
- Korean law — Korea's fair-use equivalent and its case law are not analysed here.
- Actual data — AI companies' acquisition routes and licence terms are largely undisclosed.
Sources
- TechCrunch — Sony Music, Warner sue Anthropic, alleging a 'brazen campaign' of intellectual property theft
- Daeryun — Analysis of AI copyright litigation against OpenAI, Anthropic and others
- IP Daily — Music publishers sue Anthropic over AI training, with pirated copies at issue
- Edaily — Sony and 35 others sue Anthropic over unauthorised training on lyrics and sheet music
- Axios — Sony, Warner sue Anthropic, alleging "blatant theft" of intellectual property