
Variational Autoencoder
A Variational Autoencoder is a neural network that transforms data such as images into a compact mathematical description and can generate new, similar data from this description. It forms one of the foundations of generative AI and is used, for example, for image generation and data compression.
A Variational Autoencoder is a neural network — that is, a system that learns from examples — that performs two tasks simultaneously. It learns to reduce data such as images or text to a very compact form. And it learns to generate new data from this compact form again, data that resembles what it learned but is not identical. This makes the Variational Autoencoder one of the so-called generative models: models that not only recognize but also create. It was introduced in 2013 by researchers Diederik Kingma and Max Welling and continues to influence how modern image generators are built to this day.
Why generative models need the VAE
Before the VAE, simpler autoencoders already existed. These can compress and reconstruct data, but cannot meaningfully interpolate. Interpolating means finding a middle ground between two examples. If you ask a simple autoencoder for something between a dog and a cat, you often get image noise instead of a plausible animal.
The VAE solves this problem. It forces the model to keep its internal representation space organized. Similar things then lie close together in the internal space, and the transition between two points still results in something meaningful. This makes it a genuine tool for data generation — not just data copying.
This property is practically valuable. Anyone who needs new training data, who wants to gently alter images, or who wants to train a model without having millions of real examples, can use a VAE to generate realistic-looking variations.
Encoder, latent space, and decoder
A VAE consists of two parts. The first part is called the encoder. It takes an input image — for example, a photo of a face — and computes a compact description from it. This description is called a latent vector, meaning a point in a multidimensional numerical space, the so-called latent space. This space might have 128 or 256 dimensions, while a high-resolution image has millions of pixels. The compression is thus enormous.
The crucial difference from a simple autoencoder: the VAE encoder does not compute a fixed point, but a probability distribution — specifically a mean and a spread. A point is then randomly drawn from this distribution. This sounds like an error, but it is intentional. Through this randomness, the latent space learns to be smooth and coherent.
The second part is called the decoder. It takes a point from the latent space and reconstructs an image from it. During training, the generated image is compared with the original, and both parts learn together. In the end, one can choose any point in the latent space — even one the model has never seen during training — and still obtain a realistic image.
VAEs in products and current research
Variational Autoencoders are found in many systems that appear in the news today. Image generators such as Stable Diffusion use a variant called a Latent Diffusion Model. Here, a VAE handles the pre- and post-processing: it compresses an image into the latent space, another model works there, and the decoder reconstructs the result back into an image. Without the VAE, the computation would run on the full pixels — which would be many times more expensive.
In medicine, VAEs are used to generate many synthetic datasets from a few real patient records. This is useful because real patient data is difficult to obtain for privacy reasons. In the pharmaceutical industry, VAEs help explore new molecular structures: one moves through the latent space and checks which generated molecules might have promising properties.
By now, more powerful alternatives to the VAE exist, such as diffusion models or GANs (that is, systems in which two networks train against each other). Nevertheless, the VAE is considered an important step in the history of generative AI. Anyone who wants to understand how modern image generators think cannot avoid the VAE.