Authors Guild v. Google, October 2015: Fair Use
On this page6 sections
Authors Guild v. Google, October 2015: Fair Use
- Date:
- October 16, 2015
- Location:
- New York (Second Circuit Court of Appeals)
- Paper/Outcome:
- Authors Guild, Inc. v. Google, Inc., 804 F.3d 202 (2d Cir. 2015) — fair use ruling
- Significance:
- Established fair use for book scanning; now underpins every LLM training defence
A class-action copyright lawsuit filed in 2005 against Google Books, alleging that scanning books for search constituted infringement. The Second Circuit’s 2015 fair-use ruling became the leading precedent for text-and-data-mining fair use, and the most-cited precedent in 2023-2024 AI training cases.
The Google Books Project
Authors Guild v. Google was a class-action copyright lawsuit filed in 2005 against Google Books, alleging that scanning millions of books for full-text search constituted infringement. The Second Circuit’s October 2015 fair-use ruling became the leading precedent for text-and-data-mining fair use, and the most-cited precedent in 2023-2024 AI training cases.
The Google Books project — originally called Google Print, and later renamed Google Book Search — was one of Google’s most ambitious undertakings. The idea was simple in concept but staggering in scale: scan every book ever written, make the full text searchable, and provide users with snippets of the relevant passages. The project was announced in December 2004, with five library partners: the University of Michigan, Stanford University, Harvard University, the University of Oxford (specifically the Bodleian Library), and the New York Public Library. Google would scan the books in these libraries, create digital copies, and incorporate the text into its search index. When a user searched for a term, Google would return, alongside the usual web results, relevant books — with a “snippet view” showing a small portion of the text around the search term.
The project was, from the beginning, controversial. Google was scanning books without seeking permission from the copyright holders. For books in the public domain — books whose copyrights had expired — this was unproblematic. But for books that were still under copyright, the scanning was, in the view of many authors and publishers, a clear infringement. Google’s argument was that the scanning was fair use — that the use was transformative (it was for search, not for providing the full text), that it did not substitute for the original work (a snippet view was not a substitute for reading the book), and that it served a public benefit (making the contents of millions of books searchable, for the first time, was a service to scholarship, research, and the broader public).
The Authors Guild, which represented thousands of published authors, disagreed. The Guild argued that Google’s scanning was a massive, systematic infringement of copyright — that Google was making complete digital copies of millions of copyrighted books, without permission, and that the fair use defence did not apply. The Guild filed its lawsuit in September 2005, and the case began its long journey through the federal courts.
The Settlement Attempt
Before the case went to trial, the parties attempted to settle. The proposed settlement — the Amended Settlement Agreement (ASA), submitted to the court in 2009 — was, by most accounts, one of the most ambitious and controversial proposed settlements in the history of copyright law. The ASA would have created a massive digital library — a registry of all the books Google had scanned, with mechanisms for authors to claim their works, to receive payments for the use of those works, and for Google to sell access to the full text of out-of-print but still-copyrighted books. The settlement would have, in effect, created a new legal framework for the digital use of books, and it would have given Google a dominant position in the market for digital books.
The ASA attracted enormous opposition. Authors, publishers, libraries, academics, competitors (especially Microsoft and Amazon), and governments all filed objections. The objections ranged from antitrust concerns (the settlement would give Google a monopoly on digital books — the concern that drove the Justice Department’s opposition and Judge Chin’s rejection) to privacy concerns (Google would know what books every user was reading) to class-action concerns (the settlement would bind all authors, including those who had not consented). The Justice Department investigated the antitrust implications, and it filed a statement of interest arguing that the settlement raised serious antitrust concerns.
On March 22, 2011, Judge Denny Chin — the federal judge overseeing the case — rejected the ASA. Judge Chin’s ruling held that the settlement was not “fair, adequate, and reasonable” to the class of authors it purported to represent, and he suggested that an “opt-in” alternative — where authors would have to affirmatively choose to participate, rather than being automatically included — might be more appropriate. The rejection of the ASA was, for Google and for the Authors Guild, a significant setback, and it sent the case back to litigation.
The Fair Use Ruling
With the settlement rejected, the case proceeded to summary judgment. On November 14, 2013, Judge Chin — who had been elevated to the Second Circuit in April 2010 but was still overseeing the case as a district judge by designation — granted summary judgment for Google, finding that the scanning was fair use. The ruling was, for Google, a complete victory. Judge Chin’s analysis applied the four factors of fair use, as established by Section 107 of the Copyright Act (17 U.S.C. § 107):
The first factor — the purpose and character of the use, including whether it is commercial or transformative — favoured Google. Judge Chin found that Google’s use was “highly transformative.” Google was not providing the full text of the books; it was providing a search function that allowed users to find relevant books and to see brief snippets. The use was commercial (Google is a for-profit company), but the transformative nature of the use outweighed the commercial character.
The second factor — the nature of the copyrighted work — also favoured Google, or was at least neutral. The books at issue included both fiction and nonfiction, and while fiction is entitled to more protection than nonfiction, the factor was, in the court’s view, of limited significance in this case.
The third factor — the amount and substantiality of the portion used — was the factor that most favoured the Authors Guild. Google had scanned the entire text of every book. But Judge Chin found that the full-text scanning was justified by the purpose of the use — you could not build a search index without scanning the full text. The amount used was, in this sense, appropriate to the transformative purpose.
The fourth factor — the effect of the use upon the potential market for the copyrighted work — favoured Google. Judge Chin found that Google Books did not serve as a substitute for the books. A user who found a book through Google Books and saw a snippet was not less likely to buy the book; if anything, the user was more likely to buy it, because Google Books served as a discovery tool. The Authors Guild had argued that Google Books harmed the market for digital licensing of books, but Judge Chin was not persuaded.
Taken together, the four factors favoured Google, and Judge Chin found that the scanning was fair use. The Authors Guild appealed.
The Second Circuit
The Second Circuit Court of Appeals heard the case and, on October 16, 2015, issued its ruling (804 F.3d 202). The court, in a unanimous opinion (panel: Judges Pierre N. Leval, José A. Cabranes, and Barrington D. Parker, Jr. — Chin, though a Second Circuit judge by then, was not on the panel, having authored the decision under review), affirmed Judge Chin’s ruling. The Second Circuit’s analysis was, in its essentials, the same as Judge Chin’s — the scanning was transformative, it did not serve as a market substitute (a finding central to the fourth fair-use factor), and the four factors of fair use favoured Google. The court’s opinion was, however, more detailed and more carefully reasoned than Judge Chin’s, and it provided a clearer framework for analyzing fair use in the context of large-scale digitisation.
The Second Circuit’s ruling was, for the Authors Guild, the end of the road — at least in the lower courts. The Guild petitioned the Supreme Court to hear the case, but on April 18, 2016, the Supreme Court denied certiorari, leaving the Second Circuit’s ruling intact. The denial of certiorari was, in effect, the final word: Google’s book scanning was fair use, and the decade-long legal battle was over.
The Significance
The significance of Authors Guild v. Google extends far beyond the specific case of book scanning. The ruling established, for the first time, that the large-scale digitisation of copyrighted works — without permission — for the purpose of building a search index is fair use. This principle — that making a complete digital copy of a copyrighted work, for a transformative purpose that does not substitute for the original, is fair use — has become one of the most important precedents in American copyright law.
The ruling’s significance has, in the 2020s, become even more apparent. The AI companies that are now being sued for training their models on copyrighted text — OpenAI (sued by the New York Times, see B81), Stability AI (sued by Getty Images, see B80), and others — are relying, in significant part, on the Authors Guild v. Google precedent. The argument is analogous: the AI companies are making complete digital copies of copyrighted works (text, images, video) for the purpose of training their models, and they argue that this use is transformative (the models do not reproduce the original works; they learn statistical patterns from them) and that it does not substitute for the originals (a language model is not a substitute for reading the New York Times). Whether the analogy holds — whether the courts will extend the Authors Guild v. Google reasoning to the AI training context — is one of the central legal questions of the AI era, and the answer will shape the future of the AI industry.
The Google Books project itself, as of 2026, continues to operate. Google has scanned over forty million titles, in more than five hundred languages, and the search function remains available to users around the world. The project, despite the decade of legal uncertainty, has become an indispensable tool for researchers, students, and readers — a testament to the power of the technology and to the legal framework that allowed it to survive.
The Legacy
The legacy of Authors Guild v. Google is, by any measure, the legacy of fair use in the digital age. The ruling established that fair use — a doctrine that was, at its origins, a modest exception to copyright — could be applied to large-scale, systematic digitisation, and that the transformative use of copyrighted material (using a copyrighted work for a different purpose than its original), even without permission, could be lawful. The ruling was, for the technology industry, a green light — a signal that the courts would, in appropriate circumstances, allow the use of copyrighted material for transformative purposes, and that the development of new technologies would not be blocked by copyright claims.
The legacy is also, in important ways, a contested one. The Authors Guild and the broader community of authors and creators have argued, from the beginning, that the ruling gives too much power to technology companies and too little protection to creators. The argument — that Google’s scanning, however transformative, was built on the unauthorised use of millions of authors’ work, and that the authors deserved compensation — is, in many ways, the same argument that is now being made against the AI companies. The question of whether the Authors Guild v. Google precedent will be extended to the AI context, or whether the courts will distinguish the cases, is one of the most consequential legal questions of the 2020s. The answer will determine, in significant part, who controls the data that trains the AI systems of the future, and on what terms.
- Authors Guild, Inc. v. Google, Inc., 804 F.3d 202 (2d Cir. 2015) — the Second Circuit ruling, 16 October 2015. law.justia.com/cases/federal/appellate-courts/ca2/13-4829/13-4829-2015-10-16.html
- Authors Guild, Inc. v. Google, Inc., 954 F. Supp. 2d 282 (S.D.N.Y. 2013) — Judge Chin’s district court ruling, 14 November 2013.
- 17 U.S.C. § 107 — the four-factor fair-use statute. law.cornell.edu/uscode/text/17/107
- Google Books — the project, still operational. books.google.com
- Wikipedia: Authors Guild, Inc. v. Google, Inc. — a consolidated secondary source for the case history. en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc.
- Connection to NYT v. OpenAI (B81) and Getty v. Stability AI (B80) — the live cases that are testing whether this precedent extends to AI training.
This piece is part of Minds & Machines: Beyond the Series. The companion pieces B80 — Getty Images v. Stability AI and B81 — NYT v. OpenAI & Microsoft are the live AI-copyright cases testing whether this 2015 fair-use precedent extends to AI training, and B78 — Stable Diffusion Release is the open-weights image generator whose release triggered the image-AI copyright wave.
Was authors guild v. google inevitable — the product of forces too large to redirect — or was it a series of choices, each of which could have gone differently? The answer matters, because it determines whether the future is something that happens to us or something we make.
Subscribe
Get new articles delivered to your inbox. No spam — just the story behind the screen.