SuperGPQA

SuperGPQA is a test comprising around 26,000 highly difficult expert-level questions from 285 academic disciplines, used to assess how much genuine expert knowledge an AI system truly possesses. It is regarded as the successor to the smaller GPQA test and also covers rare fields of study that are missing from typical exams.

SuperGPQA is a large collection of exam questions for computer programs that can understand and answer text. Such collections are known in technical jargon as benchmarks, meaning standardized comparison tests. The questions come from 285 different academic disciplines, ranging from medicine and mechanical engineering to library science. In total there are around 26,000 tasks, each with multiple answer options to choose from. The collection was published in early 2025 by a research group associated with the Chinese company ByteDance. The name refers to an older, much smaller test called GPQA, which SuperGPQA significantly expands upon.

Why 285 subjects instead of just physics, chemistry, and biology

The predecessor GPQA contained only a few hundred questions from three natural sciences. That is a narrow basis for judging general expert knowledge. If a system performs well there, it tells you little about whether it can also give useful answers in agricultural science or legal history. SuperGPQA closes this gap by deliberately including subjects that are rarely tested.

Behind this lies a well-known problem: test results lose their informative value when the tasks are already contained in the system’s training material. Experts call this contamination. Well-known exam questions circulate on the internet and thus easily end up in the data from which a system learns. The system can then simply reproduce the answer from memory without actually grasping the underlying matter. New and rare questions from niche subjects reduce this risk.

The level of difficulty is also important. Older knowledge tests have become too easy for today’s systems; many achieve over 90 percent correct answers on them. Such tests can no longer distinguish between good and very good systems. On SuperGPQA, the best results at the time of publication stood at about 60 percent. This leaves room for improvement, and progress becomes visible.

How the questions were created

The tasks were not generated automatically but created through a multi-stage process. First, human experts collected questions from textbooks and exam materials in their respective fields. Language models then helped standardize the wording and generate additional incorrect answer options. In the end, humans once again reviewed the result. The authors call this mixture of human and machine a human-LLM collaboration process.

A central filter was the question of whether a task was actually difficult enough. Questions that several simple systems answered correctly right away were discarded. What remained were tasks that require genuine expert knowledge or multiple reasoning steps. Each question has up to ten answer options instead of the usual four. This reduces the chance of getting it right through pure guessing from 25 to about 10 percent.

An example makes the principle tangible. Instead of asking which element carries the symbol Fe, SuperGPQA demands, for instance, the evaluation of a reaction process under certain conditions. The first is pure recall knowledge, the second requires application. It is precisely this distinction that the test aims to make visible.

SuperGPQA in model announcements and leaderboards

You will primarily encounter the name in technical reports on new AI models. When a company introduces a new system, it typically lists a table of benchmark results. Alongside names like MMLU, GPQA, or HLE, SuperGPQA now frequently appears there as well. Tech news articles also pick up on such figures when reporting on progress.

Some skepticism is warranted when interpreting such reports. A percentage figure in a table says nothing about how reliably a system performs in your specific everyday use. Benchmarks measure narrowly defined abilities under laboratory conditions. A model can excel at expert questions and still miscalculate dates or fabricate sources.

Another common misconception is that a high score means understanding in the human sense. What is measured is only how often the selected answer matches the stored solution. Why a system got it right remains an open question. SuperGPQA is therefore a useful comparative tool, but not a certificate of intelligence.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.