By Kiley Tan and Sophie Freeman
Though it’s universally acknowledged that Artificial Intelligence (AI) is transforming society, it is still difficult for many businesses to predict exactly what the impact of AI will be for them. How do they benefit from the opportunities it brings? And how do they mitigate the potential risks?
As there’s a lot to cover, we’re splitting our update into two. In the first blog, we will clarify how AI works, and give you some information about Text and Data Mining (TDM). Next month we’ll talk about the outputs of AI and how to navigate the Intellectual Property minefield that comes with it.
If you would like our specific help in your business, please contact us on 020 3056 8538 or info@thelegaldirector.co.uk.
How do large language models work?
From the outset, it is worth stating that AI is a misnomer – it is not intelligent. It does not think or reason. Generative AI is simply a large language model that uses algorithms to carry out very intensive and accurate predictive statistical analysis as to what word would follow the next. This gives the appearance of coherence, but the model is entirely reliant on “learning” what is or isn’t correct.
So, if you feed the model exclusively with bad jokes, you will get bad jokes (with a small possibility of getting a really good joke). Current models are trained by the content on the internet – web pages, government documents etc. ChatGPT (one such AI model) can be asked to review an incorrect answer, for example, so it can ‘self-correct’ and become more coherent in the next answer it provides.
What is ‘data scraping’ and ‘text and data mining’?
As the efficacy of these models requires large amounts of data, that data is usually obtained from the websites on the internet. In this area, there are two very distinct concepts:
- Data scraping is the process of extracting information from websites. Once collected, the data is structured in a useful format, but it is not processed or analysed.
- Text and data mining (TDM) focuses on the analysis of large data sets to discover patterns and gain insights.
For AI models, TDM is usually the process used.
EU/UK/US differences on TDM
Does this mean that your website can be subject to TDM by third parties?
Unfortunately, the UK, the EU and the US have different legal approaches to this problem meaning that the effect on businesses will vary depending on the territories of operation.
TDM is generally NOT permitted in the UK due to existing UK copyright law (Copyright, Designs and Patents Act 1998 or “CDPA”), the one exception to this being for TDM for non-commercial purposes e.g. research. After a campaign by the creative industries, the UK Government has decided that it will not change copyright laws further to permit wider use of TDM.
In the EU, regulated by the Digital Copyright Directive (which does not currently apply in the UK), the opposite is true. Use of material extracted or processed by TDM is a permitted exception to historical copyright restrictions, UNLESS PROHIBITED. Rightsholders can still opt out. The TDM exception may be disapplied, therefore, where the rightsholder has expressly reserved the right “in an appropriate manner such as machine-readable means”. This brings with it the additional complication of the extent to which the bots, which carry out TDM, can “understand” and respect these reservations.
In the US, the doctrine of “fair use” of copyright works has generally been viewed as favourable to text and data mining practices, but this will be put to the test in the legal challenges brought against Stability AI before the US courts.
Increasing legislation and case law across the world should ultimately clarify legitimate practices in what are currently uncharted and fairly murky waters. However, in such uncertain times, it pays to be cautious.
How do UK business restrict or prevent TDM?
TDM poses the threat of infringement of IP Rights for businesses. But protecting these rights technically, practically and legally remains difficult. TDM uses repetitive software programs (or bots) to view and extract data from websites – there is no actual person copying the data. This creates a legal issue as to whether bots would be bound by a website’s terms of use. This issue remains unclear in the UK.
IP controls and protection of web pages and online content
Cross-border issues are likely to be complex in this area.
In the UK, Copyright protection is granted by the application of all or part of the CDPA to works originating in certain other countries and in relation to certain nationals. Reciprocal protection may also be granted to works from other countries on the basis that they provide adequate copyright protection under their own laws, to protect British works and nationals.
In addition, the Digital Copyright Directive is likely to be relevant to any UK business that operates a website that could be said to be directing its activities within the EEA or any UK companies which have group companies in EU member states, due to the application of intra-EU jurisdiction.
Therefore, if a UK business directs its online activities to the EU and wants to exclude the effects of the Digital Copyright Directive, it should include an express reservation of its rights in this regard in website terms and conditions. It is also worth considering whether the location of the server “storing” a website might also be a relevant factor in assessing where any infringement takes place.
What can businesses do to protect their IP against TDM?
You may want to restrict access to your valuable content, IP-rights data and materials, by putting it behind a paywall or requiring logins or other technical authentication barriers (such as Captcha). It is also worth amending your website terms of use to explicitly state that data extraction is forbidden. To be effective in the EU, this should be in a prescribed form.
Using AI to your advantage
If you want to use TDM and AI to enhance your business, you may be considering the possibility of creating a new revenue stream by developing your own AI-generated language model. Take care – assess the proposal and requirements from a technical and legal (IP) perspective to ensure that you do not inadvertently include protected works belonging to a third-party into the model. This could potentially lead to a claim for IP rights infringement, which could derail the whole project.
Get in touch
Please get in touch on 020 3056 8538 or info@thelegaldirector.co.uk if you would like our help in your business. We can conduct an IP audit and work with you to exploit the potential of AI in your business and mitigate against IP infringements.
We also have lots of valuable IP-related information on the website. You can download Sophie’s Guide to Protecting IP here, and listen to her ‘Lunch & Learn’ on Copyright here.
Related Posts
-
In this thought piece, Client Legal Director Kiley Tan tells ICON readers what role Artificial Intelligence can play in meeting their environmental, social, and corporate governance targets.

