AI and Copyright: Who Owns the Training Data?
On this page4 sections
AI and Copyright: Who Owns the Training Data?
The legal disputes over whether training a generative AI model on copyrighted works constitutes fair use or infringement. AI copyright cases — NYT v. OpenAI, Getty v. Stability AI, Authors Guild suits — turn on whether model weights are derivative works and whether outputs that resemble training data infringe.
The Central Question
Every generative AI system — every large language model, every image generator, every music synthesiser — is trained on data. The data comes from the internet: web pages, books, articles, images, code, videos. Much of this material is copyrighted. News articles are copyrighted by the news organisations that publish them. Books are copyrighted by their authors. Photographs are copyrighted by the photographers who take them. Songs are copyrighted by the songwriters and record labels that produce them.
The AI companies — OpenAI (the San Francisco AI lab founded in December 2015, behind ChatGPT), Google, Anthropic, Meta, Stability AI, and others — collected this material, used it to train their models, and did not ask permission. They did not pay for it. They did not even tell the creators that their work was being used. They simply scraped it — collected it from the internet using automated programs called crawlers — fed it into their training pipelines, and built multi-billion-dollar businesses on top of it.
The creators are not happy about this. They argue that their work has been used without their consent, without compensation, and without credit. They argue that the AI models trained on their work compete with them — generating text that substitutes for articles, images that substitute for photographs, code that substitutes for programmers. They are suing.
The AI companies argue that what they are doing is legal. They argue that training an AI model on copyrighted material is fair use — a US copyright-law doctrine allowing use of copyrighted material without permission under certain circumstances. They argue that the models do not reproduce the training data (the billions of words or millions of images a model is trained on); they learn statistical patterns from it. They argue that the models are transformative — they create new outputs, not copies of the inputs.
The question of who is right is the most important unresolved legal question in the AI industry. The answer will determine whether AI companies can continue to train their models on internet data without paying for it, or whether they will need to license the data — a cost that could run into billions of dollars and that would give large content owners enormous leverage over the AI industry.
Fair Use: The Legal Framework
Fair use is a doctrine in United States copyright law that allows the use of copyrighted material without the copyright holder’s permission, under certain circumstances. The doctrine is designed to balance the rights of copyright holders with the public interest in the free flow of information and ideas.
Fair use is not a simple rule. It is a four-factor test, codified at 17 U.S.C. § 107, and courts weigh all four factors together:
-
The purpose and character of the use. Is the use transformative — does it add new meaning, message, or purpose to the original work? (Transformative use is the most-litigated sub-element of the first factor; the first factor also weighs whether the use is commercial or non-commercial.) Commercial uses are less likely to be fair use than non-commercial uses, but commercial use does not automatically disqualify a use from being fair.
-
The nature of the copyrighted work. Is the original work factual (like a news report) or creative (like a novel)? Factual works receive less copyright protection than creative works.
-
The amount and substantiality of the portion used. How much of the original work was used? Using a small excerpt is more likely to be fair use than using the entire work.
-
The effect of the use on the potential market. Does the use harm the market for the original work? If the AI model generates outputs that substitute for the original work, this factor weighs against fair use.
The first fair-use factor — purpose and character of the use — turns largely on whether the use is “transformative,” meaning it adds new meaning, message, or purpose to the original work rather than merely repackaging it. The Supreme Court established this test in Campbell v. Acuff-Rose (1994). In AI copyright cases, transformative use is the most-litigated issue: AI companies argue that training is transformative (the model learns statistical patterns, doesn’t reproduce the text); creators argue it is not (the model can, under certain conditions, reproduce the text verbatim, and its outputs compete with the originals in the same market). Recent rulings in 2025 (Anthropic, Meta) have turned heavily on this factor.
The AI companies’ argument is that training a model on copyrighted text is transformative — the model does not reproduce the text, it learns statistical patterns from it. They cite two key precedents: Authors Guild v. Google (2015), in which the Second Circuit Court of Appeals held on October 16, 2015 that Google’s book-scanning project was fair use, and Authors Guild v. HathiTrust (2014), which reached a similar conclusion on June 10, 2014 about a digital library’s book-scanning for search and accessibility purposes.
Authors Guild v. Google is the key precedent AI companies invoke — and the one creators distinguish. In October 2015, the Second Circuit Court of Appeals unanimously held that Google’s mass digitisation of millions of books for the Google Books project was fair use. The court emphasised that Google Books served a search and indexing function: users could find books and view short snippets, but Google did not display the full text or generate competing expressive output. AI companies argue that training a language model is analogous — the model learns patterns, it doesn’t reproduce the books. Creators argue the analogy fails: Google Books helped users find books; generative AI produces new text that competes with the books. Whether the search-vs-generation distinction is decisive remains an open legal question.
The creators’ argument is that training a model on copyrighted text is not transformative — the model can, under certain conditions, reproduce the text verbatim, and the model’s outputs compete with the original works in the same market. They argue that the Google Books precedent does not apply, because Google Books provided a search function (allowing users to find books), not a generation function (allowing users to create new text that substitutes for books).
The courts have not yet resolved this question. The NYT v. OpenAI case, the Getty v. Stability AI case, the Authors Guild v. OpenAI case, and — by 2025 — more than 60 copyright lawsuits against AI companies are all working their way through the courts. (There were roughly 30 active AI-copyright suits by the end of 2024; the number more than doubled during 2025.) The outcomes of these cases will, collectively, determine the legal framework for AI training.
The Licensing Market
While the courts are deciding, a parallel market has emerged: the licensing market — the market for AI training-data licenses. OpenAI, in particular, has been signing licensing deals with major content owners. These include:
- The Associated Press — among the first news organisations to license content to OpenAI (announced July 13, 2023)
- Axel Springer — the Berlin-based German media group (Politico, Business Insider, Bild, Die Welt), announced December 13, 2023
- News Corp — Rupert Murdoch’s media conglomerate (Wall Street Journal, New York Post, Times of London), announced May 2024, reportedly worth more than $250 million over five years
- The Financial Times (April 2024), Vox Media and The Atlantic (May 2024), Le Monde (March 2024), and Reddit (May 2024)
These deals give OpenAI the right to use the licensors’ content for training, and they serve as evidence — from the creators’ perspective — that OpenAI knows its unlicensed training was unlawful. From OpenAI’s perspective, the deals show good-faith pursuit of voluntary licenses.
The licensing market creates a strategic split among content owners. On one side: News Corp, Axel Springer, the Financial Times, Vox Media, The Atlantic, AP, Le Monde, and Reddit — all chose licensing. On the other: the New York Times, the Daily News, the Authors Guild, Getty Images, and others — all chose litigation. The split reflects different assessments of the legal landscape, different business strategies, and different views about the future relationship between content creators and AI companies. Notably, the licensors tend to be larger organisations that can negotiate from strength; the litigants include both large organisations (NYT, Getty) and collective-representation bodies (Authors Guild) acting on behalf of creators who could not negotiate individually.
Getty Images — the stock-photo agency — sued Stability AI in the UK on January 16, 2023, and in the US (District of Delaware) on February 3, 2023, over Stable Diffusion’s training data. The New York Times sued OpenAI and Microsoft on December 27, 2023, alleging that training ChatGPT on NYT articles infringed copyright.
If the courts rule that training on copyrighted material is fair use, the licensing market may collapse — if the training is legal, why would AI companies pay for licenses? If the courts rule that training on copyrighted material is not fair use, the licensing market will explode — every AI company will need to license every piece of content it uses for training, and the content owners will have enormous leverage.
What’s at Stake
The stakes are enormous. If the courts rule in favour of the AI companies, the current model — train on internet data without paying — will be validated. The AI industry will continue to grow, unencumbered by licensing costs, and the creators of the text, images, and audio that feed the models will receive nothing.
If the courts rule in favour of the creators, the AI industry will face a fundamental restructuring. AI companies will need to license their training data, which will be expensive and complex. Large content owners — the New York Times, News Corp, Getty Images, the major record labels, the major book publishers — will gain significant leverage over the AI industry. Smaller creators may benefit from collective licensing schemes, but they may also find that their individual work is not valuable enough to license. And the open-source AI movement, which relies on freely available training data, may be severely constrained.
The question of who owns the training data is not just a legal question. It is a question about the relationship between the creators of content and the builders of AI systems. It is a question about who benefits from the AI revolution, and who pays for it. And it is a question that the courts — slowly, case by case — are beginning to answer.
- “The New York Times Company v. Microsoft Corporation” — Complaint, December 27, 2023 — The most consequential AI copyright lawsuit. See the companion piece B81 for the full story.
- “Getty Images v. Stability AI” — UK High Court, filed January 16, 2023; US District of Delaware, filed February 3, 2023 — The leading image-AI copyright case. See the companion piece B80 for the full story.
- “Authors Guild v. Google” — Second Circuit Court of Appeals, October 16, 2015 (804 F.3d 202) — The key fair-use precedent that the AI companies are relying on.
- “Authors Guild v. HathiTrust” — Second Circuit Court of Appeals, June 10, 2014 (755 F.3d 87) — The parallel digital-library fair-use ruling.
- “AI Copyright Litigation Tracker” — Copyright Alliance — A regularly updated tracker of all AI copyright lawsuits (60+ as of 2025).
- “Does ChatGPT Violate New York Times Copyrights?” — Harvard Law School — A balanced academic analysis of the legal questions.
This piece is part of Minds & Machines: Beyond the Series. The main series covers the broader creative-AI story in A26 — The Future of Creativity. The companion pieces B81 — NYT v. OpenAI and B80 — Getty v. Stability AI cover the individual lawsuits in detail.
Was ai and copyright inevitable — the product of forces too large to redirect — or was it a series of choices, each of which could have gone differently? The answer matters, because it determines whether the future is something that happens to us or something we make.
Subscribe
Get new articles delivered to your inbox. No spam — just the story behind the screen.