Schematischer Ablauf eines Diffusion Transformers: Ein verrauschtes Bild wird in gleichgroße Patches aufgeteilt, die zusammen mit einem Textbefehl in den Transformer-Block eingegeben werden. Nach mehreren Entrauschungsschritten entsteht ein fertiges Bild.

Diffusion Transformer

A Diffusion Transformer is an architecture for AI image generators that combines two proven principles: the step-by-step denoising of images and the processing structure of modern language models. It forms the foundation of many current image and video generators, including Sora and Stable Diffusion 3.

When an AI system generates an image, it often starts with pure noise — that is, an image full of random pixels — and refines it step by step until a sharp result emerges. This process is called diffusion. For a long time, a particular network form was used for this, the so-called U-Net, which treats images in its structure like the layers of an onion. A Diffusion Transformer replaces this U-Net with a different underlying structure: the transformer. Transformers are originally known from text processing — they power ChatGPT, for instance. The idea behind the Diffusion Transformer is to use this proven structure for image generation as well.

Why the choice of underlying structure matters

The U-Net was the standard for diffusion models for years and works well for fixed-size images. But it has a decisive drawback: it scales poorly. This means that a larger U-Net eventually brings hardly any better results, no matter how much computing power is thrown at it.

Transformers show exactly the opposite. The bigger they get and the more data they see, the better they become — almost without an upper limit. This property is called scalability. So anyone who wants to generate better images with more computing effort is better served by a transformer as the foundation. This is the real reason the industry switched to Diffusion Transformers.

On top of that, transformers can process text and images simultaneously. This makes it easier to generate a matching image from a text prompt like “sunset over a mountain range” — the transformer understands both inputs in the same language.

From noise to image: the process within a DiT

For a transformer to work with images, the image must first be brought into a form suitable for it. Images consist of millions of pixels — too many to process directly. That’s why the model first divides the image into small, equally sized tiles, called patches. Each tile is treated as a kind of data point, similar to a word in a text. The transformer can then reason about all the tiles simultaneously.

Then the actual denoising begins. The model receives a noisy image and an instruction — for example, the text prompt — and must estimate which noise it should remove. This step is repeated many times, typically twenty to a thousand times, until a clear image emerges. Each repetition improves the result by a small amount.

The crucial mechanism in the transformer is called attention. It allows the model to relate every tile of the image to every other tile. This enables the model to capture global relationships: if there’s a sun on the left side of the image, the model knows that the sky on the right should be warmly colored.

DiTs in current products and research debates

The term was coined in 2022 by a research paper by William Peebles and Xinchen Shen, who showed that transformers can outperform U-Nets in image quality. Since then, the architecture has become the new standard. Stable Diffusion 3 and its successor use Diffusion Transformers as their core. Midjourney and Adobe Firefly also work with related approaches.

The architecture received particular attention when OpenAI introduced the video generation model Sora in February 2024. OpenAI confirmed that Sora is based on a Diffusion Transformer — specifically one that can also handle videos of varying length and resolution, because patches are formed flexibly in three-dimensional space (width, height, time).

In tech news, the term mainly comes up when new image generators are introduced or compared. A common misconception: some people confuse the Diffusion Transformer with the transformer model from text processing. Both share the same underlying structure but are trained for completely different tasks. A Diffusion Transformer doesn’t write text — it denoises images.

Related Products

Latest News

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.