Summary: The Text and Data Mining Exception under the Copyright and related rights in the Digital Single Market Directive

This summary by Julienne Te explores the Text and Data Mining (TDM) exceptions under the EU’s Digital Single Market (DSM) Directive, focusing on their relevance to AI development. It explains how copyrighted works can be used to train AI models under two key exceptions—for scientific research and for broader purposes subject to opt-out by rights holders—highlighting the legal balance between innovation and copyright protection.

Download

As technology evolves, new questions about Intellectual Property Rights (“IPR”) emerge, particularly in the development of Artificial Intelligence (“AI”). Commonplace AI models are heavily reliant on the use of works protected by IPRs for their continuous development.  

AI models typically “learn” by processing massive amounts of text and data, which are often compiled in training datasets. In many cases, these datasets comprise material protected by copyright, including books, articles, websites, blog posts, videos, and other literary and artistic works available online.

The use of these copyrighted works to train AI models raises several critical questions. At the outset, there are differing views as to whether IPRs can help or hinder innovation in the field of AI. Indeed, given that training data sets often contain protected works, how can AI models be developed without infringing copyright? On the other hand, how could authors or artists prevent their works from being used to train AI models, if they so choose?

In this respect, the Text and Data Mining (“TDM”) exceptions under the EU’s Digital Single Market (“DSM Directive”) Directive would be of relevance to AI models today.

Article 2(2) of the DSM defines text and data mining as “any automated analytical technique aimed at analysing text and data in digital form in order to generate information which includes but is not limited to patterns, trends and correlations”.

In simple terms, TDM occurs when data is analysed digitally to extract information. Common applications of this in the AI field are found in machine learning, where extensive datasets that may include copyright-protected works are processed for training AI systems.

With that said, questions concerning the permissible use of copyrighted works for TDM arise for both copyright holders and AI developers. The conditions for this use under the DSM are discussed below.

Download to continue reading