
Watermarking
Watermarking refers to techniques that embed hidden markers into texts, images, or other content to make their origin verifiable later. In AI, it is used to reliably label content generated by a model as machine-produced.
When an AI model writes a text or generates an image, the result looks at first glance just like human work. Watermarking is a technique that counteracts this: it embeds an invisible marker into the content that can later be read out like a stamp of authenticity. The watermark alters the content in a barely perceptible way, but remains detectable for a matching recognition program. The idea originates from copyright protection, where watermarks have long been used in photos and music. For AI content, it is becoming important today for a different reason: not to protect ownership, but to make deception detectable.
Why watermarks are needed for AI content
Language models and image generators produce content that is barely distinguishable from human-made content. This creates a problem: misinformation, fabricated quotes, or deceptively realistic images can be produced en masse. If no one can tell whether a text originates from a human or a machine, it becomes difficult to assign responsibility.
Watermarking is meant to close exactly this gap. An operator deploying an AI model could be required to mark all generated content with a watermark. Authorities, editorial offices, or platforms could then automatically check whether a suspicious text is of machine origin. The European Union has stipulated in the AI Act that AI-generated content must be labeled — watermarking is considered one of the possible technical ways to achieve this.
How a watermark is embedded in a text
For images, a watermark can be described as a fine, invisible alteration of individual pixels. For text, this is more difficult, because letters do not have continuous values that could be minimally shifted. A common approach therefore works via the choice of words itself.
When a language model generates a text, at each point it selects the next word from several plausible candidates. In watermarking, these candidates are divided in advance into two groups — for example, “green” and “red” words — and the model preferentially chooses from the green group. For the reader, this makes no difference: the text sounds normal. But anyone who knows which words belong to the green group can statistically check whether a text contains a conspicuously high number of them. If that is the case, the watermark is detectable — even if the text appears inconspicuous at first glance.
An important difference from a regular stamp is that the watermark is embedded in the generation process itself, not added afterward. It also cannot simply be “removed” without substantially altering the text.
Where watermarking appears in practice today
Google has developed a watermarking technology called SynthID for its image generator Imagen and has also introduced it for texts from the Gemini model. OpenAI has announced it will integrate similar methods into ChatGPT. Meta is likewise researching automatically labeling AI-generated content on Facebook and Instagram.
In the school context, the topic is especially debated: could teachers use a watermark detector to check whether an essay was written by an AI? In principle, yes — but only if the model used embedded a watermark and the detector matches it. Anyone using a model without watermarking, or heavily rephrasing the text afterward, leaves no trace.
This is precisely where the central weakness lies: watermarking only works if all relevant providers participate. A single major provider without a watermark is enough to undermine the protective effect for the entire ecosystem. That is why technical solutions and regulatory obligations usually accompany this topic together.