Direct Preference Optimization

Direct Preference Optimization

Direct Preference Optimization (DPO) is a method used to train an AI language model to better satisfy human preferences – without requiring a separate reward model for this purpose. It is considered a simpler and more stable alternative to RLHF, the approach that was standard until then.

Once a language model has finished training, it can form sentences – but it cannot yet reliably judge which answer is more helpful or correct. To improve this, it is shown examples of human judgments: here are two answers to the same question, which one is better? From these comparisons, the model learns what people prefer. Direct Preference Optimization, DPO for short, is a method that does exactly this. It has been widely used since 2023 and has replaced an older, more elaborate approach called RLHF – Reinforcement Learning from Human Feedback – in many projects.

DPO’s role in the fine-tuning of language models

A model that has only been trained on raw text material answers technically correctly – but often not in the way a human would wish. It may respond too long, too vaguely, or in the wrong tone. DPO is the fine-tuning step that changes this. It takes place at the end of the training process and shapes the model’s behavior in practice.

Without this step, chatbots like ChatGPT or the assistant on a smartphone would hardly be usable in everyday life. The base knowledge is present after pretraining, but only preference optimization makes answers polite, clear, and to the point. Many of the quality differences users perceive between various AI products arise precisely at this step.

Pairwise comparisons instead of point-based judges

The training material for DPO consists of triples: a question, a preferred answer, and a rejected answer. Humans – or sometimes another AI model – have rated these pairs. DPO uses these comparisons directly to adjust the model’s weights. It increases the probability that the model chooses the preferred answer and decreases it for the rejected one.

The crucial difference from RLHF lies in the architecture. RLHF first trains its own reward model, which learns to predict human judgments, and then uses this as an arbiter. This is a two-stage process involving two models, which is error-prone and computationally intensive. DPO skips the intermediate stage entirely. It mathematically derives which model adjustment best fits the comparisons and applies it directly.

A well-known problem with RLHF is so-called reward hacking: the model learns to trick the arbiter instead of actually becoming better – similar to a student who optimizes for a teacher’s quirks instead of the subject matter. DPO avoids this risk because there is no separate arbiter to trick.

DPO in products and research

DPO was introduced in a 2023 research paper from Stanford University and spread quickly afterward. Most modern open-source models – such as the Llama or Mistral families – use DPO or closely related variants for their fine-tuning. Commercial providers, too, have adopted the method or are experimenting with it.

In the trade press, DPO mainly appears in reports about new model versions. When a company announces that its model is “better aligned” or “safer in handling sensitive topics,” DPO or a similar method is often behind it. The term “alignment” – the question of whether a model does what people actually want – is closely linked to DPO.

Several further developments of DPO now exist, such as IPO or SimPO, which aim to fix individual weaknesses of the original method. The basic idea remains the same: pairwise comparisons instead of elaborate intermediate models. DPO has thus sparked an entire research direction concerned with the question of how to translate human judgments into a model as directly as possible.

Latest News

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.