GRPO

GRPO

GRPO is a training method that improves AI models through reward signals – without requiring a second, resource-intensive auxiliary model. It became widely known through its use in DeepSeek's R1 model and is considered a particularly resource-efficient alternative to older methods.

An AI model is first trained to predict text. This makes it linguistically fluent, but not yet particularly useful or reliable. To specifically train it toward good answers, another phase is needed: the model receives feedback on which answers are better and which are worse – similar to school grades. GRPO, short for Group Relative Policy Optimization, is a specific method for exactly this phase. It belongs to the family of so-called reinforcement learning methods, i.e. learning procedures that work with rewards and penalties. What’s special about GRPO is how it calculates these evaluations: not in absolute terms, but always in comparison to a group of answers to the same question.

Why GRPO makes the difference

Older methods from the same family – such as PPO, which was long considered the standard – require a second, almost equally large auxiliary model in addition to the actual model. This so-called critic model continuously estimates how good an answer was. This costs enormous amounts of computing power and memory.

GRPO manages without this auxiliary model. Instead, it calculates the evaluation directly from the group of compared answers. This sounds like a technical detail, but it has considerable consequences: models for which PPO would be too expensive can only be trained at all with GRPO. The method opens up elaborate training methods even for smaller research groups and companies that don’t operate huge data centers.

The group principle behind GRPO

The name reveals how it works: for a single question, the model generates several answers simultaneously – a group. Each answer receives an evaluation from an external source, the so-called reward model or a rule-based check. Then GRPO calculates which answers within this group were better than average and which were worse.

The model is then adjusted so that it produces above-average answers more frequently in the future. An important difference from simpler approaches: the model always compares itself with itself, not with a fixed target point. This means the training automatically adapts to the current performance level. A model that is already good has to hold its own against a high bar.

To keep the training stable, GRPO also prevents the model from changing too quickly and too drastically. There is a kind of distance rule: each training step may only shift the model’s behavior up to a defined degree. This protects against the model overwriting previously learned knowledge just because a single training question causes a strong deviation.

GRPO in practice – DeepSeek and its consequences

GRPO became known in early 2025, when the Chinese AI lab DeepSeek announced that it had used the method for its model DeepSeek-R1. R1 achieved results on mathematical and logical tasks that could keep pace with significantly more expensive models. Since DeepSeek simultaneously published its program code, other research groups were able to directly replicate GRPO and adopt it into their own projects.

Since then, GRPO has regularly appeared in specialist articles and tech news when it comes to so-called reasoning models – that is, models that don’t simply output an answer, but think through a solution path in several steps. GRPO is particularly well suited for this purpose because the quality of such steps can be well verified with rules: a mathematical derivation is either correct or incorrect, and no human judgment is needed for every single line.

For the broader development of AI, GRPO is an example of the fact that the decisive progress sometimes doesn’t come from more computing power, but from a clever simplification of an existing method. Anyone reading about the training costs of large language models will come across GRPO as one of the reasons why these costs have recently dropped significantly.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.