
Encoder-only model
An encoder-only model is a type of AI language model that reads and understands text but does not generate new text itself. It is used primarily when a model needs to grasp the meaning of a sentence – for example, to search documents or detect sentiment in texts.
An encoder-only model is a language model that fully reads in a text and translates it into a kind of internal representation — a list of numbers. These numbers describe what the text means. In doing so, the model does not itself produce output as readable text. Instead, it delivers a compressed representation of meaning that other systems can work with further. That sounds unspectacular, but it is exactly the right approach for many practical tasks.
Strengths in text understanding
An encoder-only model reads every sentence in both directions at once. So it sees not only what lies to the left of a word, but also what comes after it to the right. This makes it better at capturing the meaning of a word within its context.
The word “bank” means something different in “He sits on the river bank” than in “He transfers money at the bank.” An encoder-only model can recognize this difference because it processes the entire sentence at once. A model that only reads from left to right, on the other hand, has to guess before it knows the end of the sentence.
That is precisely why encoder-only models are often better than generative models — which produce text word by word — at comprehension tasks. Anyone who doesn’t want to write text but wants to classify text doesn’t need a generator.
Processing from left and right simultaneously
The best-known encoder-only model is called BERT, developed by Google in 2018. It is first pretrained on a huge amount of text — in the process, it learns to fill in missing words in sentences. This step is the actual language understanding: the model has to understand what fits in a sentence before it can fill the gap.
Afterward, the model is specialized for a specific task, for example classifying sentences as positive or negative. This second step is called fine-tuning. The base model remains the same; only the final building block is swapped out and retrained. This saves the effort of training from scratch for every new task.
Technically, at its core lies what is known as a Transformer — an architecture that calculates relationships between all the words in a sentence simultaneously. Encoder-only models use only the first half of this architecture: the part that reads in and understands. They omit the part that outputs text — hence the name.
Use in search, analysis, and classification
Encoder-only models are embedded in many everyday applications without being visible. Search engines use them to match the meaning of a search query with the meaning of web pages — not just by keywords, but by semantic similarity. Google has used BERT for the majority of its search queries since 2019.
Sentiment analysis on social networks, the automatic sorting of customer reviews, or the detection of spam in emails also frequently runs on encoder-only models. Wherever a system has to decide what a text means or which category it belongs to, these models are the first choice.
In the news, the term mainly comes up in comparison with large generative models like GPT. The key difference: GPT and similar systems generate text, while BERT and its successors understand it. For many enterprise applications, the smaller, specialized encoder-only models are faster, cheaper, and more accurate than a universal text generator.