Rasch model

Rasch model

The Rasch model is a statistical method from test theory. From correct and incorrect answers, it simultaneously estimates how difficult each task is and how able each person is – on a shared scale.

If someone solves 20 out of 30 tasks on a test, that number alone says little. Were the tasks easy or hard? And how do you compare two people who worked on completely different tasks? The Rasch model is a computational method that solves exactly this problem. It assumes that the probability of a correct answer depends on only two quantities: the ability of the person and the difficulty of the task. Both quantities are estimated from the collected answers of many people and plotted on the same scale. The model is named after the Danish mathematician Georg Rasch, who developed it in the 1950s.

Why raw scores alone are not comparable

The simplest way to evaluate a test is to count the correct answers. However, this raw score always depends on which tasks happened to be included in the test. A student with 18 points on a hard test may be stronger than one with 24 points on an easy test. Raw scores therefore cannot be compared across different test versions.

The Rasch model cleanly separates the two influences from each other. A person’s estimated ability remains theoretically the same regardless of which selection of tasks they worked on. Conversely, a task’s estimated difficulty remains the same regardless of which group worked on it. Experts call this property specific objectivity. It is the reason why large school studies such as PISA work with such models.

In practice, this means that different cohorts or countries can receive different task booklets and still appear on a common scale. The prerequisite is that some tasks overlap. These shared tasks act like bridge pillars between the test booklets.

Ability and difficulty on one scale

In the Rasch model, each person receives a value and each task receives a value on the same axis. What matters alone is the distance between the two. If ability and difficulty are equal, the chance of a correct answer is exactly 50 percent. If the person is clearly more able than the task is difficult, the chance rises toward 100 percent. If they are clearly weaker, it falls toward zero.

One way to picture this is a high jump bar. The bar has a height, the jumper has jumping power. Whether the attempt succeeds depends only on the difference, not on the absolute height. The Rasch model, however, calculates with probabilities rather than a certain outcome, because people sometimes have a good day and sometimes a bad one.

The relationship between distance and probability follows a fixed S-shaped curve, the logistic function. A computer tries out different values for people and tasks until the predicted response patterns match the real ones as closely as possible. An important restriction is this: all tasks must have the same curve shape. Related models with more parameters allow tasks to differ in how well they distinguish between strong and weak test-takers. The Rasch model deliberately forgoes this and thereby gains its clean comparison properties.

From PISA to AI benchmarks

The method is best known from school achievement studies. The PISA scores around 500 points are not counted tasks but converted Rasch estimates. Language tests, medical questionnaires, and admission tests are also frequently evaluated this way. In psychology, the model is also used to check whether a questionnaire actually measures a single trait at all.

An important application is adaptive computer-based testing. After each answer, the program re-estimates ability and then selects a task whose difficulty fits as well as possible. This means significantly fewer questions are needed for an equally accurate result.

More recently, the model has also appeared in AI research. Language models are tested with collections of tasks known as benchmarks. Instead of merely reporting the percentage of correct answers, researchers use Rasch models to estimate the model’s ability and the difficulty of each test task. This makes results more comparable across different task collections. One common misconception, incidentally, is that the model delivers absolute truth. It delivers estimates, and these are only as good as the assumption that a single ability is truly being measured.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.