Elo Rating

Elo Rating

The Elo rating is a score that calculates a participant's playing strength based on wins and losses in duels. Originally invented for chess, it is now also used to rank AI models according to which answer humans prefer in direct comparisons.

The Elo rating is a number that expresses how strong someone is in direct comparison with others. It was invented by the physicist Arpad Elo for chess. Every player starts with a score, usually around 1500. After each game, a certain amount of points moves from the loser to the winner. What matters is not just who wins, but against whom. A win against a much stronger opponent brings many points, a win against a weak one almost none.

Why a single number says more than a win rate

The simplest way to measure strength would be the share of games won. But this number is easy to manipulate. Anyone who only plays against weak opponents quickly reaches 90 percent wins. Anyone who always faces the best looks worse on paper. The Elo rating solves this problem because it factors in the strength of the opponent.

The second advantage is that the number allows for a prediction. From the difference between two Elo ratings, the expected probability of winning can be derived. A 100-point lead means roughly a 64 percent chance of winning, a 400-point lead about 91 percent. An Elo rating is therefore not just a ranking, but a statement about how a duel is likely to turn out.

This is precisely why the AI industry has adopted the system. The question of which language model gives the better answer can hardly be settled with a test. But you can show people two answers and ask which one they prefer. Out of many such duels, an Elo rating emerges for each model.

How points change hands

Before each duel, the system calculates an expected outcome. To do this, it compares the scores of the two participants. Afterwards, it looks at what actually happened. The difference between expectation and reality determines how many points shift.

An example makes this tangible. A model with 1200 points faces one with 1400. A win by the stronger one is expected. If the weaker one wins anyway, the expectation was clearly wrong, and many points change hands. If the stronger one wins, it only gets a few, since that was expected anyway.

How large the shifts turn out to be is controlled by a factor called the K-factor. A high K-factor lets the number react quickly to new results, but also makes it jittery. A low one keeps values calm, but reacts sluggishly to genuine progress. It is also important to note: the Elo rating has no absolute meaning. 1500 points only mean something in comparison to the other participants, not on their own.

Elo leaderboards for chatbots and games

The Elo rating remains best known in chess. There, the world elite sits above 2800 points. Many online games also use similar systems to match up fair opponents. Anyone who is paired with a similarly strong opponent in matchmaking owes this to such scores.

In the AI world, the term is mostly encountered on comparison platforms like LMArena, formerly Chatbot Arena. There, users ask a question and receive two answers from two anonymous models. They pick the better one, and only afterwards are the names revealed. Out of millions of such votes, a leaderboard with Elo-like scores emerges.

In news about new models, this value regularly appears as evidence. Still, it should be read with caution. It measures what people spontaneously prefer, not necessarily what is factually correct. Politely phrased, well-structured answers often score better, even if they contain an error. The Elo rating is therefore a useful indicator, but not proof of quality.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.