Psychometric Test

A psychometric test is a standardized procedure that expresses traits such as intelligence, knowledge, or personality in numerical form. In the AI world, such tests are now also applied to language models in order to make their capabilities comparable.

Some things can be measured directly: height with a tape measure, weight with a scale. Traits like concentration, logical thinking, or anxiety, on the other hand, cannot be touched. A psychometric test is a procedure that translates such invisible traits into a number anyway. To do this, it presents many people with exactly the same tasks under exactly the same conditions. A value is calculated from the answers, which can be compared with the values of other people. The best-known case is the intelligence test, but employment aptitude tests at companies and many questionnaires in psychology also belong to this category.

Why a number says more than a gut feeling

People are notoriously inaccurate at judging other people. Someone who finds another person likeable often automatically considers them more competent as well. A standardized test is meant to eliminate exactly this effect, because all participants receive the same tasks. That is why companies, schools, and clinics have relied on such procedures for decades.

A good test must meet three requirements. It must be reliable, meaning it delivers roughly the same result upon repetition. It must be valid, meaning it truly measures what it claims to measure. And it must be objective, meaning it functions independently of who administers and evaluates it. If any one of these three conditions is missing, the neat number at the end is worthless.

This is precisely where the biggest criticism lies. A test can be formally sound and still disadvantage people, for instance when the tasks require prior knowledge from a specific cultural background. Anyone who is unfamiliar with the phrasing of a question fails because of the language, not because of their thinking ability. Numbers therefore appear more objective than they often actually are.

From answer to score

At the beginning there is always a collection of tasks, called items in technical language. These items are first tested on a large group, often several thousand people. A comparison scale, the so-called norm, emerges from their results. An individual result only gains meaning through this comparison.

The well-known IQ score is a good example of this calculation logic. The average of the comparison group is set to 100 by definition. A score of 115 does not mean that someone solved 115 tasks. It means that the person is clearly above the average of their age group. The number describes a position in comparison, not an absolute quantity.

Tasks that everyone answers correctly or everyone answers incorrectly are removed again during development. They do not differentiate between participants and therefore provide no information. In the end, a test preferably consists of items of medium difficulty. This explains why such tests often seem oddly thrown together.

Psychometrics in language models and the application process

In everyday life, one encounters psychometric tests mainly during job applications. Many corporations have candidates solve online tasks on logic, numerical reasoning, and work style. The personality questionnaires found on the internet also imitate this principle. Serious procedures differ from fun tests in that their quality has been scientifically verified and published.

In tech news, the term has been appearing in a new context for several years now. Researchers have language models work through classic intelligence and personality tests and report the scores. Exams for humans, such as bar or medical exams, are also used as benchmarks for AI. Such results serve as selling points when a company introduces a new model.

This transfer is controversial, however. The norm values originate from humans, and a model may have already seen the test tasks in its training data. In that case, the test measures memory rather than thinking ability. Anyone reading reports about a chatbot's IQ should therefore carefully check which test was used and where the comparison figures come from.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.