Ablaufskizze eines Text-zu-Bild-Modells: links der eingegebene Text, der über einen Text-Encoder in eine Zahlendarstellung umgewandelt wird; rechts daneben eine Reihe von fünf Bildkacheln, die von reinem Zufallsrauschen über zunehmend erkennbare Formen bis zum fertigen Bild eines Fuchses im Schnee führen; Pfeile zeigen, dass die Textdarstellung jeden Entrauschungsschritt steuert.

Text-to-Image Model

A text-to-image model is a computer program that generates a matching image from a written description. Well-known examples are Midjourney, Stable Diffusion, and DALL·E.

A text-to-image model is a program that turns a written sentence into an image. You type in, for example, “a red fox in the snow, oil painting” and after a few seconds you get a matching image. This image did not exist before. It is not pulled from a database, but calculated pixel by pixel from scratch. That’s why two images for the same sentence never look exactly alike. The best-known programs of this kind are called Midjourney, Stable Diffusion, and DALL·E.

What has changed for photographers, designers, and forgers

Until a few years ago, a usable advertising image cost money and time. You needed a camera, a studio, or an illustrator. Today, a sentence and a few cents of computing costs are often enough. This affects entire professions: stock photo agencies, illustrators, graphic design departments. The company Getty Images has therefore sued a provider.

The dispute revolves mainly around the training material. The models learned from billions of images from the internet, often without the permission of the creators. Artists argue that their work is embedded, unpaid, in these programs. The providers claim that no image is copied, but rather that only a style is learned. Courts in several countries are only now clarifying this.

On top of that, there is a problem of misuse. Realistic images of events that never happened can be produced within seconds. Fake photos of politicians or disasters spread quickly on social networks. The EU therefore requires, in its AI Act, that such images be labeled as artificially generated.

From noise to image: the diffusion process

Most of these programs work using a method called diffusion. During training, real photos are taken and image noise is gradually poured over them until only static remains. The model learns to reverse this process. It has to estimate, from a noisy image, what the clean image underneath looked like. It repeats this task millions of times.

When generating, the program then starts with pure random noise. Step by step, it removes part of it, usually over twenty to fifty passes. Gradually, shapes, colors, and details emerge. A sculptor who chisels away everything from a block of stone that doesn’t belong to the figure is a fitting comparison. Except the model doesn’t start with stone, but with color noise.

The typed-in sentence steers this process. A second network converts the text into a long sequence of numbers representing its meaning. At every denoising step, the model checks whether the emerging image matches this sequence of numbers. That’s how you end up with a fox and not a dog. This is also exactly where the most common weakness lies: numbers, captions, and hands are often rendered incorrectly, because the model imitates shapes rather than counting.

Where these images show up today

They are most visible in advertising and on news sites. Many article images for abstract topics are now artificially generated. Product designs, book covers, game graphics, and presentation slides are also created this way. In online shops, retailers test generated variants of the same product photo.

You also encounter the technology as a built-in feature. Photoshop can fill in missing parts of an image, phone apps replace the background of a photo. ChatGPT and similar assistants generate images directly in the chat window. Users often don’t even notice that a dedicated model is behind it.

In stock market news, the term usually comes up in connection with Nvidia, Adobe, or legal disputes. A related term is the text-to-video model, which generates moving images using the same principle. A common misconception, by the way, is that these programs search the internet. They are offline and don’t contain a single stored photo, only learned patterns.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.