
GPQA-Diamond
GPQA-Diamond is a collection of 198 particularly hard science exam questions used to test how well an AI program can reason. The questions are designed so that they can barely be solved even with an internet search.
GPQA-Diamond is a test for computer programs that can understand and answer text. It consists of 198 questions from biology, chemistry, and physics. All questions are multiple-choice: there are four possible answers, exactly one of which is correct. They were written by people who are pursuing or have already completed a doctorate in their field. The special thing about it: even trained experts from a different field usually fail at these questions, even though they are allowed to search the internet for as long as they like. That’s why the G in the name stands for “Google-proof.”
Why 198 questions alone carry so much weight
Companies like OpenAI, Google, or Anthropic regularly claim that their new model is smarter than the previous one. Such claims need evidence. A test with fixed questions and one fixed correct answer produces a number that can be compared. GPQA-Diamond is one of these numbers and appears in almost every product announcement.
The test was published in 2023 because older exams had become too easy. On many school and university tasks, the programs already scored above 90 percent. A test on which everyone can do almost everything no longer distinguishes anything. When it was released, GPQA-Diamond was so hard that the best systems of the time stayed below 40 percent.
By now, the strongest models reach over 85 percent. That sounds like a triumph, but it is above all a sign that this test, too, will soon have run its course. Experts call this benchmark saturation: when there’s hardly any room left at the top, a two-percentage-point difference says little. With only 198 questions, one percentage point corresponds to roughly two tasks.
How the questions were created and vetted
The original GPQA comprises 448 questions. Diamond is the hardest selection drawn from it. A question only made it into this selection under two conditions. First, experts from the same field had to solve it correctly in the majority of cases. Second, experts from an unrelated field had to fail at it, despite having internet access and no time limit.
This dual test is the core of the matter. The first condition ensures that the question actually has a clear-cut solution and is not simply poorly framed. The second condition ensures that looking things up alone is not enough. A question that can be answered with three clicks measures research skills rather than understanding.
A common misconception is the assumption that the model could simply have memorized the answers. This cannot be ruled out entirely, since the questions are freely available on the internet. That is precisely why experts pay attention to whether a model shows its solution path. If a program arrives at the answer through several traceable intermediate steps, that speaks more for genuine reasoning than for recall.
Reading GPQA-Diamond in product announcements
When a company presents a new language model, it almost always shows a table with test results. GPQA-Diamond usually appears there alongside math tests and coding tasks. The score is considered an indicator of how well a model handles multi-step technical reasoning. It says nothing about how well the model writes, translates, or organizes appointments.
When reading such tables, it’s worth taking a look at the fine print. Some results are produced by having the model answer the same question multiple times and counting the most frequent answer. Others give the model plenty of computing time to think. Both inflate the score and make comparison with a single simple run unfair.
The lower bound also matters. With four possible answers, pure guessing yields about 25 percent. A result of 30 percent is thus barely better than chance. And a high score does not mean one should blindly trust the model in a chemistry class: a test made up of 198 questions only covers a small slice of an entire field.