HLE-Verified

HLE-Verified

HLE-Verified is a vetted subset of the demanding knowledge benchmark "Humanity's Last Exam," in which experts have double-checked every question and every model answer. It is meant to prevent AI systems from failing on flawed questions and thereby appearing worse than they actually are.

To compare how good different AI systems are, they are made to sit standardized exams. Such a collection of test questions is called a benchmark. One of the hardest of these tests is called “Humanity’s Last Exam,” or HLE for short. It consists of several thousand questions from mathematics, physics, medicine, linguistics, and many other fields, submitted by experts from all over the world. With so many submissions, errors inevitably creep in: incorrectly stated solutions, ambiguous wording, or questions that simply cannot be answered from the material provided. HLE-Verified refers to the portion of these questions that independent reviewers have re-checked and found to be clean.

Why a single wrong model answer skews the entire leaderboard

A benchmark is only as reliable as its answer key. If the stored answer to a question is wrong, a model that calculated correctly gets penalized. Conversely, a model can score points simply because it happens to make the same mistake as the question’s author. In an easy test this barely matters. But with HLE, results initially sat in the range of just a few percent.

That’s exactly the core of the problem. If a system solves 8 percent of the questions while roughly 10 percent of the questions are themselves flawed, the measurement contains more noise than signal. The gap between two competing models then becomes meaningless. Yet companies advertise precisely such gaps, and investors read them as proof of progress.

The vetted subset is meant to remove this noise. Scores measured on HLE-Verified therefore tend to come out somewhat higher than on the full test. This is not a sudden leap in capability. It simply means that unfair deductions no longer apply.

How the questions go through the review process

The process starts with the originally submitted questions. First, automated procedures pre-sort obvious problem cases: duplicate questions, broken formulas, unclear formatting. After that, at least one person with subject-matter expertise in the relevant field examines the question. They check three things: Is the question unambiguous? Is the given solution correct? And can it actually be derived from the question at all?

Questions on which the reviewers cannot agree are dropped when in doubt. This is a deliberate decision. Better a smaller, clean test than a large one with unclear cases. The remaining questions form the verified set, which is then used for official comparisons.

An important distinction from related terms: HLE-Verified says nothing about whether an AI model can justify its answer. It is exclusively about the quality of the exam questions, not the quality of the answers. Anyone who confuses “verified” with “verified AI” is mistaken. It is the test that was verified, not the test-taker.

Where the figure shows up in announcements

When a company unveils a new language model, a table of benchmark results is almost always included. HLE appears there as a particularly demanding entry, often with the addition “verified” or a footnote about the subset used. News sites then pick up these figures in their coverage of the rivalry among the major AI providers.

When reading such reports, it’s worth taking a close look. Was the measurement taken on the full HLE or on the verified subset? Was the model allowed to search the internet or use tools while doing so? And how many attempts did it get per question? Two percentage figures are only comparable if these conditions match.

The term thus stands in for a larger issue: benchmarks are themselves products that get maintained, corrected, and eventually replaced. The closer systems get to the maximum score, the less the test actually reveals. At that point a new one is needed — and the cycle starts all over again.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.