Schema: Ein Sprachmodell wählt das nächste Wort aus einer Liste, die ein geheimer Schlüssel in grüne und rote Wörter teilt. Das Modell bevorzugt grüne Wörter. Der fertige Text geht an ein Prüfprogramm, das mit demselben Schlüssel den Anteil grüner Wörter zählt und daraus auf maschinelle Herkunft schließt.

Watermark for AI Texts

A watermark for AI texts is a hidden pattern that a computer program deliberately embeds while writing a text. Anyone who holds the matching key can later prove with high probability that the text originated from a machine.

Programs like ChatGPT write texts that are barely distinguishable from those written by humans. A watermark is meant to make this distinction possible anyway. In doing so, the program embeds a hidden pattern into the text while writing. For readers, the text remains completely normal. But anyone who knows what to look for can extract the pattern again using a verification program. The name comes from the watermark on banknotes: it’s there, you just need to look closely to see it.

The fight against text floods and forgeries

Texts from machines now cost almost nothing to produce. This suddenly makes certain problems very large. False reports can be reworded a hundred times over within minutes. Advertising copy floods search engines. And in schools, it’s unclear who actually wrote a given piece of homework.

A watermark promises a technical way out here. It’s supposed to provide a reliable answer to the question: human or machine? That is exactly what lawmakers are demanding too. The European Union’s AI Act requires that artificially generated content be labeled in a machine-readable way. Providers of large AI systems therefore have to deal with the issue, whether they want to or not.

But there’s also a self-interest on the part of the companies. If AI-generated texts flow into the internet en masse, later models end up training on the output of their predecessors. Experts call this model collapse. Quality declines across generations as a result. A watermark helps to filter out such texts beforehand.

How the pattern gets into word choice

A language model writes word by word. At each step, it calculates a probability for thousands of possible next words. Usually there are several fitting continuations. After “The weather today is”, words like “nice”, “good”, “warm”, or “lousy” would all make sense. It’s precisely within this freedom that the watermark hides.

The most common method randomly splits the vocabulary into two halves before each step. One half is called the green list, the other the red list. Which words are green is determined by a secret key together with the preceding word. The model then slightly favors the green words. It doesn’t force anything—it just nudges a little.

For a single word, this means nothing. A human would, purely by chance, hit roughly half green words. But across a hundred words, a clear imbalance emerges. A verification program with the same key simply counts the green hits. If the proportion is well above half, the text very likely came from the machine. This method requires no archive of every text ever generated.

The weak point lies in rewriting. Anyone who runs the text through a second AI program or swaps out many words destroys the pattern bit by bit. Short texts are unreliable anyway, since there aren’t enough words for statistics. A watermark is therefore an indication, not proof.

Images are already marked, texts hardly at all

For images and videos, labeling is already commonplace. Google marks images from its models using the SynthID method. Many cameras and programs support the C2PA standard, which attaches provenance data to the file. Images have millions of pixels in which a pattern can be hidden well.

With texts, the situation is different. OpenAI developed a working method but did not release it for years. The reason is commercial: users might switch to a competitor without a watermark. Google, at least, has released SynthID for texts as open source.

A common misconception concerns the detectors used by schools and universities. These tools guess based on stylistic features and don’t know any key. They are frequently wrong and particularly disadvantage people writing in a foreign language. A genuine watermark works completely differently, but it can only be used if the provider builds it in in the first place.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.