
Decoder-only model
A decoder-only model is a design of language model that generates text word by word from left to right, always looking back only at what has already been written. Most well-known chatbots, including ChatGPT and Gemini, are based on exactly this principle.
A decoder-only model is a particular way of building a language model. It generates text word by word — always from left to right. In doing so, it only ever looks back at what has already been written, never at what is still to come. This principle sounds simple, but it underlies most of today’s large language models: GPT-4, Gemini, LLaMA and many others all follow this blueprint. The name comes from the fact that older models consisted of two parts — an encoder and a decoder — and this variant does away with the encoder entirely.
Why this blueprint dominates the AI industry
Leaving out the encoder sounds like a loss. In practice, it has proven to be a strength. A decoder-only model can handle a surprising number of tasks with a single architecture: translating, summarizing, writing code, answering questions. There’s no need to build it anew for every task.
The second reason is training. Decoder-only models learn according to a particularly simple principle: they see the beginning of a text and have to predict the next word. This is simple to formulate, but when exposed to enormous amounts of text, it produces models with broad knowledge. This training procedure is also called “autoregressive prediction” — auto, because the model always uses its own previous outputs as input.
An encoder-decoder model like the older Google Translate was built for translation and excelled at it. Decoder-only models are somewhat less specialized for that, but universally applicable. In an industry that bets on a single jack-of-all-trades rather than many specialists, this has turned out to be a decisive advantage.
How a decoder-only model generates text
The model is given a starting text, the so-called input or “prompt”. It then calculates which word is most likely to come next — and appends it. Then it calculates the word after that, this time based on the now-longer text. This process repeats until a punctuation mark or a set limit stops the model.
So that the model can keep track of the context of the entire text so far, it uses a mechanism called attention. It allows the model, when generating each new word, to “attend” to all previous words — that is, to establish connections even when two related words are far apart. The crucial restriction, however, is this: the model only ever looks backward. Future words do not yet exist at the time of generation and thus remain invisible. This restriction is called “causal masking” and is what distinguishes a decoder from an encoder.
How much text the model can keep track of at once is called the context or the “context window”. A short window means the model forgets earlier parts of a long conversation. Modern decoder-only models have context windows of hundreds of thousands of words — a technical advance that was unthinkable just a few years ago.
Decoder-only models in products and headlines
Anyone using ChatGPT, Microsoft Copilot, or the AI assistant on a smartphone is almost always using a decoder-only model. This blueprint also works behind the scenes in many search engines and coding aids such as GitHub Copilot. OpenAI’s GPT series made the term well known — the “G” in GPT stands for “Generative”, the “P” for “Pre-trained”, the “T” for “Transformer”, the underlying computational technique.
In tech news, the term often appears in comparisons. There, decoder-only models are sometimes contrasted with so-called encoder-only models like BERT, which don’t generate text but instead understand and classify it — for example, for search algorithms or spam filters. The boundary between the two worlds is increasingly blurring, as newer models try to combine the strengths of both camps.
A common misconception in reporting: size and blueprint are often conflated. Not every large model is automatically a decoder-only model, and not every decoder-only model is necessarily large. There are lightweight variants that run on a smartphone, and there are decoder-only models with several hundred billion parameters — the unit of measure for the number of learnable values in the model — that fill entire data centers.