L

Lord Tarassenko (CB)

Speaking in the House of Lords on 12 May 2025

Debate

Data (Use and Access) Bill [HL]

Contribution

My Lords, I rise to speak as the founder of two AI spin-outs, and I draw the House’s attention to my registered interests as the founder-director of Oxehealth, a University of Oxford spin-out that uses AI for healthcare applications. I am also the author of three copyrighted books. Since these amendments were last debated in the House of Lords, there has been a lot of high-profile comment but very few attempts, if any, to bring AI developers and creators together in the same room. During the same period, however, more businesses from the creative industries and the publishing sector have agreed content-licensing deals. That is because access to curated, high-quality content to fine-tune large language models—the step after pre-training which provides high-accuracy responses—is increasingly being monetised. Even the Guardian Media Group, a strong supporter of the creative industries, announced in February a strategic partnership with Open AI to ensure compensation for the use of its high-quality journalism. This shows that it is possible, without any change in the law, for the creative industries and the big tech companies to come to licensing agreements. The main technological development since our last debate has been the demonstration that training LLMs no longer requires the massive computer facilities and huge data centres of the big tech companies in the US. Since the beginning of the year, the Chinese company DeepSeek has released open-source LLMs hundreds of times smaller than hyperscale models such as GPT-4, Gemini or Claude Sonnet. These models, typically with, say, 10 billion weights, have been developed through the process of distillation, and they achieve almost the same level of performance as the hyperscale models with 1 trillion weights. Why is that important? It means that users of LLMs no longer have to send queries to those hyperscale models which are then processed by OpenAI, Google or Anthropic using their huge compute facilities with thousands of GPUs in their data centres. Instead, any AI developer can now train and run distilled versions of those LLMs locally on their laptops. DeepSeek was the first AI company to show how powerful the process of distillation is in the context of LLMs. Other big tech companies are now jumping on the bandwagon. In early March, Google released a brand-new LLM called Gemma 3, a lightweight, state-of-the-art open-source model that can be run anywhere from smartphones to laptops, and has the ability to handle text, images, and short videos. These open-source distilled LLMs are now being used by thousands of AI developers, in the UK and elsewhere, who are training and fine-tuning them using content, some of which may be copyrighted, publicly available on the web. Training an LLM on a laptop using data from the open web will become as commonplace as searching the web. This is already happening both within computer science departments in UK universities and in the rich ecosystem of AI start-ups and university spin-outs in the UK. A survey of 500 developers and investors in the UK AI ecosystem, carried out by JL Partners last month, had 94% of them reporting that their work relied on AI models built using publicly available data from the web, and 66% reported that if the data laws in the UK were more restrictive than elsewhere, projects would move to other countries. We need to consider the impact on the UK’s AI industry of these transparency provisions, and of the requirement to provide copyright owners with information regarding the text and data used in the pre-training, training and fine-tuning of general-purpose AI. The use of content from behind paywalls or from pirated databases such as Books3 or LibGen, which is known to have been done by Meta to train its LLM, is clearly illegal. However, for data publicly available on the open web, I would like to do a simple thought experiment to show that the transparency requirements in Motion 49A are at present unworkable. In the UK, unlike in the US, there is no copyright database. Usually, the copyright rests with the author of the work, but there are exceptions, such as when a work is created by an employee in the course of their job, and copyright may also be assigned or transferred to a third party. If we assume, generously, that it might take just one second, on average, to ascertain the copyright status of an article, book, image, or audio or video recording, on the web, it would require 31 years and eight months to check the copyright status of the 1 billion data points in a typical LLM training set—never mind thinking about setting up licensing deals with the millions of rights holders. For the distilled models that are now, as I explained, being trained or fine-tuned by UK developers, which are 100 times smaller, the copyright status check would still require one-third of a year—still an entirely unworkable proposition.

More from Lord Tarassenko (CB)

Other recent Hansard contributions by the same speaker.

About Hansard

Hansard is the official verbatim record of proceedings in the UK Parliament. Every word spoken in the Commons and Lords is recorded and published — this page is a single contribution from that record.

Partner sites