Skip to content
TEN Brief Ten verified stories a day 2026.09.01 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 3 min read · Reference

What fair use means for AI training data — why bought books count and downloaded ones do not

Fair use is the doctrine in United States copyright law that allows limited use of a protected work without permission from the rights holder. It is not a list of permitted activities but a four-factor test courts apply case by case: the purpose and character of the use, the nature of the work, the amount used, and the effect on the market for the original. AI training split along an unexpected line. In 2026 US courts accepted that training a model on lawfully purchased books can qualify as fair use, on the reasoning that training extracts statistical patterns rather than reproducing text and is therefore transformative. But acquiring the same books as pirated copies was held to be infringement, and that holding led to a large settlement. The dividing line is not what the model learned but how the copy was obtained. The music publishers' suit filed against Anthropic on August 28, 2026 targets exactly that line

A sunlit reading room at midday, two neat stacks of hardcover books on a long wooden table with an open book between them

The three lines

  • Definition — an exception allowing use without permission. Not a list, but a four-factor court test
  • The line — training on purchased books can be fair use; acquiring pirated copies was infringement
  • The point — what decided it was not what the model learned but how the copy was obtained

Key questions

What is fair use?
**An exception in United States copyright law that permits using a protected work without permission.** The most common misunderstanding first: fair use is **not a list of situations that are allowed.** It is a **test** a court applies, weighing four factors. ① **Purpose and character of the use** — is it commercial, and is it **transformative**, meaning does it create a new meaning distinct from the original? ② **Nature of the work** — factual records receive thinner protection than creative works. ③ **Amount and substantiality used** — how much, and was it the heart of the work? ④ **Effect on the market** — does the use substitute for the original or its licensing market? In practice **factors ① and ④ are usually decisive.** And a caution: Korean copyright law has its own equivalent provision, but with different requirements and a different body of case law — **US rulings do not transfer directly.**
Is AI training fair use or not?
**Training alone may qualify; acquiring the data is judged separately.** US court reasoning in 2026 split the question in two. On **training itself**, the transformative argument was accepted: a model does not store sentences to retrieve later but **extracts statistical patterns** across enormous amounts of text, and the result does not substitute for the original work. But **how the books were obtained** is a different question entirely. Downloading pirated copies is itself a reproduction-right violation, and **whatever came afterwards does not undo it.** Courts treated that portion as infringement, and a large settlement followed. Put simply: **if you bought the book and trained on it, there is an argument to have; if you torrented it, you lose before training even begins.**
So what should an AI company do?
**On the evidence so far, the direction is: be able to prove where the data came from.** The rulings turned on acquisition, not method. Three practical routes exist. ① **Purchase and licence** — costly, but the least exposed. This is why large licensing deals in news, music and images have been increasing. ② **Public-domain and permissively licensed data** — legally safest, but limited in volume. ③ **Provenance records** — keeping a durable account of what was obtained, from where, and on what terms. That is the starting point of any defence. But **far more is unsettled than settled.** Only a handful of decisions exist, appeals remain, and whether the same reasoning extends to music, video or images has not been tested. This page does **not** claim the doctrine is settled.

In the AI training-data cases, what decided the outcome was not what the model learned. It was how the book was obtained.

To see why that sentence holds, start with fair use.

1. Fair use is not a list

The most common misunderstanding, first. Fair use is not a list of "these situations are fine." It is a test a court runs.

FactorWhat it examinesPractical weight
① Purpose and characterCommercial? Transformative?High
② Nature of the workFactual record or creative workMedium
③ Amount and substantialityHow much, and was it the heart?Medium
④ Market effectDoes it substitute for the original?High

Factors ① and ④ usually decide the case. And the word transformative in ① sits at the centre of the AI debate — does the use create a different purpose and different meaning from the original?

Korean copyright law has its own equivalent provision, but with different requirements and a different body of case law. US rulings do not transfer directly.

2. Where AI training split

US court reasoning in 2026 cut a single dispute into two pieces.

StageConductHolding
AcquisitionBooks obtained by purchaseNo problem
AcquisitionBooks obtained as pirated copiesInfringement
TrainingTraining on the acquired textFair use available

The reasoning behind ② is this. A model does not store sentences and retrieve them later. It extracts statistical patterns across enormous amounts of text, and the resulting product does not stand in for the book — that is, the use is transformative and the market-substitution effect is weak.

But ① is a different question. Downloading a pirated copy is itself a reproduction-right violation. What happened afterwards does not undo it.

Hence the summary: bought books count, downloaded ones do not.

3. That line is now being aimed at

On August 28, 2026, Sony Music Publishing, Warner Chappell Music and other publishers sued Anthropic ("Sony, Warner and other publishers sue Anthropic (August 28, 2026)").

One word recurs in the complaint: torrent.

How the plaintiffs framed itWhat it targets
"Millions of pirated books via torrent"Acquisition — the side that already lost
"Lyrics and sheet music were inside"Extends the reasoning to music
CEO named personallyAims at wilfulness

The plaintiffs are not trying to relitigate whether training is transformative. That argument has already been lost once. They are placing the case on the side of a boundary that has already been drawn in their favour.

4. So what are companies doing

The direction the rulings point in is singular: be able to prove where the data came from.

ApproachStrengthLimit
Purchase and licenceLeast legal exposureExpensive
Public domain / permissive licencesLow riskNot enough volume
Provenance recordsThe starting point of any defenceHard to apply retroactively

This is the background to the growth of large licensing deals in news, music and images. The arithmetic has begun to favour paying for a licence over paying for litigation.

5. Far more is unsettled than settled

This page does not claim the above is settled doctrine. Here is what is open.

  • There are only a handful of decisions. Most are trial-level, and appeals remain.
  • A different medium may produce a different result. A holding about books does not automatically govern music, video or images. Music in particular raises the reuse of short phrases in ways text does not.
  • Outputs are a separate question. Even where training is lawful, a model producing something closely resembling the original is its own problem.
  • Jurisdictions differ. The EU has an explicit text-and-data-mining exception; Korea has yet another framework.
  • Company data sources are mostly undisclosed. Very few firms publish what they trained on.

6. What is still unresolved

  • The opinion — The "purchased is fair use, pirated is infringement" holding was not checked against the original text.
  • The settlement — Its amount, and whether an appeal is pending, were not confirmed.
  • Music — No case law exists on whether the same reasoning applies to compositions.
  • Korean law — Korea's fair-use equivalent and its case law are not analysed here.
  • Actual data — AI companies' acquisition routes and licence terms are largely undisclosed.

Sources

  1. TechCrunch — Sony Music, Warner sue Anthropic, alleging a 'brazen campaign' of intellectual property theft
  2. Daeryun — Analysis of AI copyright litigation against OpenAI, Anthropic and others
  3. IP Daily — Music publishers sue Anthropic over AI training, with pirated copies at issue
  4. Edaily — Sony and 35 others sue Anthropic over unauthorised training on lyrics and sheet music
  5. Axios — Sony, Warner sue Anthropic, alleging "blatant theft" of intellectual property

Verification

Published
Last modified
Cross-check
Checked against 5 independent sources.
Unverified
  • The 'purchased books are fair use, pirated copies are infringement' holding was not checked against the original opinion
  • The settlement amount in that case and whether an appeal is pending were not confirmed
  • No case law exists on whether the same reasoning extends to music, video or images
  • Korean copyright law's fair-use equivalent and its case law are not analysed on this page
  • The actual data provenance and licensing arrangements of AI companies are largely undisclosed
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related