
Copyright
Copyright is the right to decide who may copy, publish, or alter a work one has created. In the AI debate it is central because models are trained on vast amounts of other people's texts, images, and music.
Anyone who writes a text, takes a photo, or composes a song automatically holds rights to it. These rights are called copyright, in German usually 'Urheberrecht'. They mean: others may not simply copy, sell, or rework the work. Anyone who wants to do so anyway needs permission, often in the form of a license. A license is a contract that precisely specifies what is allowed and what it costs. The right arises without registration and without a fee, simply through the act of creating the work.
Why AI companies are ending up in court over this
Large language models and image generators learn from datasets that are almost unimaginable in scale: billions of web pages, books, photos, and songs. A very large portion of this material is protected. In many cases, the companies gathered it from the internet without asking the creators. This is exactly why a wave of lawsuits has been rolling in since 2023.
The New York Times sued OpenAI and Microsoft because their models were trained on its articles. Authors, photo agencies, and record labels have followed suit. Getty Images took action against the image generator Stable Diffusion. The sums in dispute are high; Anthropic, for instance, agreed in 2025 to a settlement of 1.5 billion dollars with book authors.
For investors, this is more than just a legal dispute. If training data has to be paid for in the future, the costs for every new model will rise significantly. At the same time, rights holdings become more valuable. Publishers, music labels, and image archives are now negotiating licensing deals worth millions instead of merely suing.
Point of contention: Learning or copying?
A model does not store texts the way a hard drive does. During training, it adjusts millions of numerical values that capture patterns in the data. AI companies therefore argue that they are not copying anything, but rather letting a system learn – similar to a human who reads many novels. In the US, they invoke the fair use rule, which permits certain uses without permission if something clearly new results from them.
The opposing side counters with two points. First, the material must first be downloaded and stored in order to be used for training, and that is technically a copy. Second, models sometimes reproduce long passages almost word for word. Experts call this memorization. It is rare, but a strong argument in court.
In the EU, a different logic applies. Text and data mining is generally permitted, but rights holders may object. This objection is called an opt-out and is stated, for example, in a website’s robots.txt file. The EU’s AI Act also requires providers to disclose, in broad terms, which data they used.
Copyright in everyday life with AI tools
The question also concerns the output, not just the training. In the US, the Copyright Office has clarified: an image generated solely from a text prompt has no human author and therefore no protection. Anyone may reuse it. As soon as a human shapes it in a recognizable way, protection can arise again for that portion.
In practice, this means for students and professionals: AI images for a presentation are usually unproblematic. It becomes riskier when the result too closely resembles a well-known work. A song in the style of a real singer, or a character that is clearly Mickey Mouse, can infringe others' rights – regardless of which program generated it.
That is why Microsoft, Google, and Adobe offer their business customers a kind of liability guarantee. They cover the legal costs if someone is sued over an AI-generated result. These commitments are a good indicator of just how uncertain the legal situation still is. Final clarity will only come once the major cases have been decided.