Humanity's Last Exam

Humanity's Last Exam

Humanity's Last Exam is a collection of roughly 2,500 extremely difficult exam questions from many fields, used to measure the performance of AI systems. The test was released in 2025 because older exams had become too easy for modern programs.

Humanity’s Last Exam is a test for computer programs that can understand and answer texts. It consists of roughly 2,500 questions submitted by experts from more than a hundred countries. The questions come from mathematics, physics, chemistry, medicine, linguistics, history, and many other fields. They are so difficult that even professors can usually only solve them within their own specialty. The name alludes to the idea that this is meant to be the last exam humans have to write for machines. The collection was published in early 2025 by the organization Center for AI Safety together with the company Scale AI.

Why the old exams had become obsolete

To compare how good different AI programs are, you need standardized tests. Such standardized tests are called benchmarks. For years, MMLU was the most important one: a collection of multiple-choice questions from school and university curricula. When this test appeared in 2020, the best programs scored around 30 percent correct answers.

A few years later, top models exceeded 90 percent. This rendered the test useless. When nearly all test-takers get almost everything right, the score no longer reveals meaningful differences. Experts call this the saturation of a benchmark.

This is precisely the gap that Humanity’s Last Exam is meant to close. The questions were deliberately chosen so that the best programs at the time would fail on them. At launch, results were below ten percent correct answers. This created renewed room at the top against which progress can be measured.

How the questions were created and reviewed

The organizers called on experts worldwide to submit difficult questions, and paid rewards for accepted contributions. Around a thousand people from about five hundred institutions took part. Roughly 70,000 questions were submitted, of which about 2,500 remained. Each question had to clear two hurdles.

First hurdle: several leading AI programs had to fail on the question. Anyone submitting a question could immediately have it tested against these models. Only what the machines could not solve moved on to the next round. Second hurdle: other experts checked the question for factual accuracy and for having a clear, unambiguous solution.

The answers are either short exact figures or multiple-choice options. This allows automatic checking of whether an answer is correct. About one in ten tasks additionally includes an image, for example a diagram or a drawing. A portion of the questions remains secret so that companies cannot deliberately train their models on them.

What the percentages in the news mean

When a company unveils a new language model, the press release almost always includes a table of benchmark results. Since 2025, Humanity’s Last Exam has been among the most frequently cited entries in it. Reports such as “Model X achieves 25 percent on HLE” refer to this test. Scores have risen significantly since launch, especially for systems that perform longer chains of reasoning before answering.

Such figures should be read with caution. Results depend on whether the model was allowed to search the internet and how much compute time it was given. A common misconception is also that a high score means general intelligence. The test measures specialized knowledge and reasoning on narrowly defined questions with a clear-cut solution.

Everyday tasks such as planning a trip or detecting a lie do not appear in it. For this reason, this benchmark too will eventually be replaced. Still, for investors and industry observers it is useful: it shows in a single number how quickly the capabilities of these systems are changing.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.