Fair Use

Fair Use

Fair use is a rule in US law that allows other people's texts, images, or music to be used in certain cases without permission. When it comes to training AI systems, it is fiercely disputed whether this exception also applies to the mass mining of other people's works.

Whoever writes a book, takes a photo, or composes a song generally decides for themselves who may copy the work. This right protects the creator. In the United States, there is an important exception to this: fair use, literally “reasonable use.” It allows other people’s works to be used in certain situations without permission and without payment. Typical cases include a quote in a review, a parody, or an excerpt used in school lessons. However, fair use is not a fixed list of permitted actions, but rather a balancing test that a court carries out in the event of a dispute.

The dispute over AI training data

Large language and image models learn from enormous amounts of data. These include newspaper articles, novels, program code, photos, and illustrations. Only a small portion of this is explicitly released for use. AI companies argue that mining this data is fair use. Their model, they say, does not copy the works but rather learns statistical patterns from them.

Publishers, authors, photographers, and record labels see it differently. They say their work was turned into the basis of a commercial product without compensation. Since 2023, numerous lawsuits have been underway in the US, including one brought by the New York Times against OpenAI and Microsoft. Some proceedings have ended in settlements and licensing agreements, while others remain open. The rulings will help determine how expensive building an AI model will be in the future.

For investors, this is no side issue. If training data must be licensed in the future, costs will rise sharply. At the same time, companies that own large data archives of their own are gaining in value. That’s why fair use rulings regularly appear in business news.

The four factors courts examine

US courts examine fair use based on four criteria. First: how is the work used? A use that creates something new has better chances than a mere copy. Such use is called “transformative.” Commercial intent tends to weigh against fair use, but does not rule it out.

Second, the nature of the work matters. Factual texts are less strongly protected than novels or poems. Third, the extent counts: a short sentence from a book carries less weight than entire chapters. Fourth, and usually decisive, the court asks about the market. Does the use take customers away from the original?

This fourth point in particular is the sore spot when it comes to AI. A chatbot that summarizes news can replace a visit to the news site. If an image model generates works in the style of a living illustrator, she may lose commissions. Importantly, all four factors are weighed together, not checked off individually. That’s why it is rarely possible to say in advance with certainty whether a use is fair use.

Why different rules apply in Germany

Fair use is a concept of US law. In Germany and the EU, this open-ended balancing test does not exist. Instead, the law lists individual permitted cases, such as the right to quote or use for teaching and research. For AI training, there is a specific rule on so-called text and data mining, meaning the automated analysis of large amounts of data. Rights holders may explicitly prohibit this use, usually via a notice on their website.

In everyday life, you mainly encounter fair use on YouTube and social networks. Reaction videos, let’s plays, and memes often invoke it. Whether it actually holds up depends on the individual case, and platforms tend to err on the side of blocking content when in doubt. A common misconception is that citing a source automatically establishes fair use. That’s not true: naming the source still doesn’t grant permission.

In the news, you’ll mostly encounter the term in connection with lawsuits against AI companies. Pay attention to whether the case concerns the training itself or the outputs of the model. Legally, these are two different questions. And a US ruling does not automatically apply in Europe, though it does shape the debate worldwide.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.