EU copyright law contains a text and data mining (TDM) exception that, among other possible uses, provides scope for AI developers to use others’ content in the training of their AI models, unless rightsholders explicitly exercise an opt‑out in a machine‑readable format. Both content creators and AI developers have complained, however, that there are shortcomings with existing opt-out mechanisms.
In a bid to find a new solution that works for both, the European Commission commissioned a study into the feasibility of a registry of TDM opt-outs. That study, which consisted of a mixture of desk research with stakeholder interviews, a survey and a multistakeholder workshop, found that current TDM opt-out mechanisms are “frequently fragmented, unevenly implemented and, in many cases, insufficiently effective in practice” and said AI developers that rely on them also “face heterogeneous and sometimes contradictory signals in automated pipelines operating at the web scale”.
The study found, though, that a new “EU-level registry leveraging digital fingerprinting technologies could be opportune and useful in supporting the effective expression of the opt-out”.
In particular, the study found that “content-based digital fingerprinting … can prove particularly useful where metadata is missing or unreliable”. It cited international standards developed for that technology: the International Standard Content Code (ISCC).
“Fingerprinting provides a content-derived handle that remains stable across file-name changes, format conversions and near-identical copies, enabling work-level identification regardless of where or how a work is shared,” the study found.
“Unlike traditional identifiers such as ISBN or DOI, which identify specific manifestations, ISCC generates a fingerprint derived from the content itself, allowing the same work to be recognised across different platforms, formats and publications. This is particularly valuable in AI training data contexts, where the scraping is done on the open web and works frequently circulate without accompanying metadata, or where rightsholders might not be able to express an optout at a sufficiently granular level,” it said.
The study acknowledged that while a new registry might support rightsholders exercise their opt-out rights effectively, it would not address the full panoply of issues arising in the AI and copyright debate. For instance, the registry would not serve as “a repository of digital assets”, nor “feature a licensing marketplace” or “adjudicate ownership”. It would sit alongside, rather than replace, “sectoral identifiers, protocols or established workflows”, the study report said.
Intellectual property law expert Gill Dennis of Pinsent Masons said a new registry has the potential to make the TDM opt-out “far more practically useful in the real world”, as it would help tackle some of the “technical barriers that currently make opt-outs difficult to identify and act on at scale”.
“At the moment, opt-out signalling relies on a patchwork of different tools, and opt-out notices can become disconnected from content as it is copied and shared online, making them harder for AI developers to find and act on,” said Dennis, who added that the registry’s value would “ultimately depend on the underlying technology and whether that can actually deliver the interoperability and cross-platform recognition that existing tools have struggled to achieve”.
Cerys Wyn Davies of Pinsent Masons, who specialises in AI and data law, highlighted how the idea of an opt-out registry is not new and said whether and how the EU study would be moved forward is likely to be watched with interest by businesses and policymakers globally.
“In December 2024, in its AI and copyright consultation paper, the UK government stated that it favoured the introduction of a TDM exception with rights reservation such that AI developers could train models on copyright works to which they had lawful access,” Wyn Davies said. “However, under those plans, rightsholders would be entitled to reserve their rights and thereby prevent their content being used for AI training by effective, accessible, machine-readable formats to prevent AI developers’ web tools accessing their content for AI training purposes. One of the technical mechanisms considered was a ‘do not train’ registry of rightsholders’ works.”
“However, the responses to the consultation raised, amongst other concerns, that the burden on creators, in particular individuals and SMEs, requiring them to opt-out would be great and impossible at scale across multiple platforms and copies, particularly as existing opt-out tools were considered insufficient. The UK government also raised the concern that existing systems lacked standardisation, accessibility and granularity. This proposal has therefore now been shelved by the UK government in favour of further evidence-gathering, avoiding changes to copyright law until it is confident reforms will meet its objectives,” she said.
“This study will be very valuable for the EU and other countries such as the UK in considering the current technical feasibility of establishing such registries for the expression of TDM opt-outs and their likely adoption and effectiveness. Following the Commission’s publication of this study, it will be important to hear the opinions of both rightsholders and AI developers on the likely workability and impact such a registry would have in practice on their respective economic and other interests,” Wyn Davies added.