Copyright Infringement & AI: A Case Study of Authors Guild v. OpenAI and Microsoft

To understand the potential legal challenges presented by AI training data in Europe, this paper, written by Maria Alexandra Mǎrginean, will analyse the Authors Guild v. OpenAI dispute through the lens of EU law. Specifically, the analysis will focus on potential copyright infringement at the input level.

Download

1.    Introduction

 

Just five days after its release in late 2022, ChatGPT-3.5 achieved 1 million users.[1] Since then OpenAI—the U.S. based artificial intelligence (AI) research and deployment company that developed ChatGPT—has experienced exponential growth, situating the company amongst the most rapidly expanding technology companies in history.[2] The impressive success of the creator of ChatGPT is based on the research and development of generative AI models. Such models are known for creating content or data similar to that which is used to train them. The produced output is based on the input (the prompt[3]) of a user.[4] Outputs can be e.g., code (like OpenAI Codex), images (such as, DALL-E, Midjourney, Stable Diffusion, DeviantArt), audio (for example, AudioCraft, Soundraw), synthetic data (MOSTLY.AI, MDClone), or text (like OpenAI GPT (Generative Pre-trained Transformer), BloombergGPT, LLaMa, Bard).

ChatGPT is a large language model (LLM) which is trained to produce human-like text.[5] To generate new content mimicking human creativity, this AI algorithm must be trained on large data sets.[6] The content produced are known as outputs. ChatGPT outputs are often plausible, however over half of its answers are false or untruthful. This is because ChatGPT was “trained to predict the next word on a large dataset of Internet text, rather than to safely perform the language task that the user wants.”[7]      

Following its success obtained by having the fastest-growing user base, questions began to circulate regarding the type of data used to train generative AI models. Such questions arose because the training sets used to train LLMs, including ChatGPT, consist of a wide range of internet text, books, and other sources. In particular, concerns were raised that these models were infringing copyright by using copyrighted works in their training material without the owner’s authorisation. These questions were initially the subject of policy, industry and academic debate but eventually found their way into the courts.

One of these cases is Authors Guild v OpenAI,[8] filed at the end of September 2023 before the court for the Southern District Court of New York. The plaintiffs are 17 fiction writers together with the Authors Guild — a professional organization for published writers with over 14,000 members. At the heart of the case are claims of copyright infringement, with OpenAI being accused of using copyrighted works without permission for the training of its GPT models.

To understand the potential legal challenges presented by AI training data in Europe, this paper will analyse the Authors Guild v. OpenAI dispute through the lens of EU law. Specifically, the analysis will focus on potential copyright infringement at the input level. Following this introduction, Chapter 2 delves into a factual analysis of the Authors Guild v OpenAI case, specifically the claims made by the plaintiffs in the US. Chapter 3 presents a short analysis of US law, providing a brief examination of the fair use doctrine. This doctrine is relevant since it permits use of copyrighted material without requiring permission from the rights owners if specific circumstances are met. The analysis of the US law serves to understand how US courts handle copyright questions regarding AI, and offers a parallel for the subsequent  analysis of EU law. To address the lack of clear legal precedent in the EU regarding text and data mining and copyright, Chapter 4 explores relevant EU legislation. Chapter 4 will focus on legislation pertaining to the right of reproduction, moral rights, and how the text and data mining exceptions in these legal instruments might apply in the context of AI training data. These exceptions allow for the legitimate use of copyrighted material in support of technological progress, research, and the digital economy. Chapter 5 provides some conclusions.


[1] ‘100+ Incredible ChatGPT Statistics & Facts in 2024 | Notta’ <https://www.notta.ai/en/blog/chatgpt-statistics> accessed 29 February 2024.

[2] ‘OpenAI on Track to Hit $2bn Revenue Milestone as Growth Rockets’ <https://www.ft.com/content/81ac0e78-5b9b-43c2-b135-d11c47480119> accessed 29 February 2024.

[3] Prompts are instructions given to an LLM for generating content. See Daniel Gervais, ‘Humans as Prompt Engineers’ (Kluwer Copyright Blog, 14 June 2023) <https://copyrightblog.kluweriplaw.com/2023/06/14/humans-as-prompt-engineers/> accessed 10 January 2024.

[4] Andrés Guadamuz, ‘A Scanner Darkly: Copyright Liability and Exceptions in Artificial Intelligence Inputs and Outputs’ (26 February 2023), p. 5 <https://papers.ssrn.com/abstract=4371204> accessed 8 November 2023.

[5] ‘What Is ChatGPT? | OpenAI Help Center’ <https://help.openai.com/en/articles/6783457-what-is-chatgpt> accessed 8 December 2023.

[6] Datasets are defined as collections of data, which are used in training and testing the algorithms. See ‘Datasets Definition | Encord’ <https://encord.com/glossary/datasets-definition/> accessed 11 March 2024.

[7] ‘Aligning Language Models to Follow Instructions’ <https://openai.com/research/instruction-following> accessed 8 December 2023.

[8] Authors Guild v OpenAI Inc SDNY 1:23-cv-08292.

This is the preview text...

Download to continue reading