This essay was first published in February 2025 as a guest post on ZDnet.com.
No data, no AI. The formula is that simple – and at the same time, so complex in its implementation. The triumphant advance of AI technologies is rolling through all industries and service sectors at a breakneck pace, fundamentally changing existing business models while simultaneously opening up unforeseen new opportunities. But immensely increased computing power or new AI models alone do not lead to the goal. It is the data that makes the crucial difference. According to Elon Musk and Ilya Sutskever, the former Chief Scientist of OpenAI, publicly accessible data sources have long been exhausted. But the demand of AI is greater – and above all, the need for high-quality, unique data.
Many organizations have extensive and valuable data sets without fully exploiting their potential. This data, collected and maintained over years, represents a highly sought-after resource – especially for AI systems, whose performance depends crucially on the quality of the training data. Data licensing agreements for sharing data for AI training have therefore become an emerging business field.
US companies such as Yelp, Shutterstock, and Reddit have already concluded corresponding agreements to license their data for training AI models. Reddit’s licensing agreement with Google, for example, has an estimated value of well over $60 million per year, while Shutterstock has struck deals with Meta and OpenAI that generate more than $100 million in revenue from generative AI alone.
In addition to text data, image, video, and audio data are particularly indispensable for AI training. Visual AI is becoming a central economic sector, provided the corresponding data sources are available. Alongside Shutterstock, Freepik is also capitalizing on this and earning well: according to media reports, around 200 million images have already been licensed. Two to four cents flow per image, leading to licensing sums of around six million US dollars. Further deals are in the pipeline. Photobucket is also in negotiations over photos and videos, with prices ranging from five cents to one US dollar per medium. These deals are not just a trend, but a fundamental shift in how companies monetize their intellectual property and data, opening up new, highly lucrative revenue streams.
In Germany, too, the topic of data sharing via licensing agreements for AI training is increasingly gaining momentum. Companies like Deutsche Bahn and the Schwarz Group, together with partners like Aleph Alpha, are developing DataHub Europe, a platform designed to promote the legally secure and efficient exchange of data. The goal is to set European standards for data sovereignty and usage while simultaneously strengthening innovations in the AI sector. The desire for “AI made in Europe” is meant to become a reality with this platform by providing data for AI models, as the thirst for high-quality data for AI training is virtually unquenchable. Especially for companies active in the European market, DataHub Europe offers a secure way to use data legally and profitably in connection with AI.
Traditional media and content providers like the news organization Reuters are also already actively navigating the licensing market. The “Reuters News” division was able to increase its revenue by 22 million US dollars. News archives, in particular, are crucial for the further expansion and professionalization of AI language models. Associated Press's contract with OpenAI includes access to articles dating back to 1985. Historical data, in particular, strengthens AI models. Axel Springer SE has also licensed historical and current content to OpenAI – in a multi-million dollar deal. Springer is thus one of the first major European publishing houses to actively participate in the AI data market, demonstrating that European media companies can tap into new business models.
Another example of European companies positioning themselves strategically in the area of data sharing is the French news agency AFP, which is also licensing its content for AI training. These developments show that Europe is not only setting regulatory frameworks but is also actively participating in the economic utilization of data.
The list of examples could go on. And they show one thing: Data – whether images, texts, or historical archives – represents a central asset for AI companies. Licensing agreements generate new revenue on the one hand, but they also raise questions regarding data protection and copyright.
As promising as the opportunities of data sharing for AI training are, they come with significant obligations. The legally secure exchange of data requires both senders and recipients to adhere to clear rules. Companies sharing data must ensure that they comply with applicable laws such as the DSGVO (German GDPR) and international regulations like the GDPR in the EU. This is not just about protecting personal data, but also about intellectual property and compliance with national regulations in the target countries.
On the recipient side, there are additional obligations. The use of data, especially for AI purposes, is subject to strict regulations like the AI Act, which sets clear guidelines on transparency and purpose limitation. Companies receiving data must ensure that the use of this data is legally safeguarded and that no ethical or legal boundaries are crossed. Transparency and compliance are the foundations for trust in the data economy.
On the one hand, the massive data hunger of AI has shifted the balance of power in favor of data providers. They now hold the upper hand – a welcome development for companies opening up new revenue streams through licensing models. But with this responsibility also grows the duty to act ethically and legally. Sustainable success in data sharing will only be possible if both sides – providers and users – take equal responsibility and stick to clear rules of the game. The opportunities are enormous, but success depends on a balanced and legally secure approach.