
GDPval
GDPval is a test published by OpenAI that checks how well AI systems perform real work tasks from paid occupations. Instead of answering exam questions, the models have to deliver things that a human would also hand in on the job – such as a legal draft, a spreadsheet, or a construction drawing.
GDPval is a test meant to measure how useful computer programs with artificial intelligence are at real professional work. The company OpenAI introduced it in September 2025. The name alludes to gross domestic product, i.e. the total economic output of a country: what’s tested are activities from industries that make up a large share of that output. The tasks don’t come from textbooks, but from experienced professionals who on average have worked in their field for many years. A program is given, for example, real documents and is supposed to produce a report, a calculation, or a plan from them. Afterwards, other professionals compare the result with what a human delivered.
Why school grades are no longer enough for chatbots
Most well-known tests for AI systems consist of questions with a clearly correct answer. Math problems, multiple-choice questions from medical or law exams, small programming tasks. Such tests are practical because they can be graded automatically. But the best models now achieve nearly full marks on them. A test that almost everyone passes with 95 percent barely says anything about differences anymore.
On top of that comes a second problem. Exam questions only remotely resemble real work. Nobody gets paid for ticking multiple-choice boxes. What gets paid is results: a finished presentation, a reviewed contract, a care plan. This is precisely the gap GDPval is meant to close. The test doesn’t ask whether a model knows something, but whether it produces something ready to hand in.
For investors and companies this is interesting because it makes the debate about automation more concrete. Instead of speculating in general terms about which jobs AI will take over, GDPval delivers figures for individual activities. However, the results should be read with caution. A single solved task is not an entire profession, and OpenAI is also testing its own products here.
Real tasks, real evaluators
GDPval comprises around 1,320 tasks from 44 occupations. These occupations come from nine economic sectors that together account for a particularly large share of the United States' economic output – including healthcare, finance, trade, administration, and manufacturing. Each task was designed by a professional from the respective occupation and provided with a model solution. On average, a task costs a human several hours of work.
The evaluation is conducted as a blind comparison. An evaluator from the same industry is presented with two results, one from the model and one from the human, without knowing which is which. They decide which one is better, or declare both equally good. The most important metric is the so-called win rate: the proportion of cases in which the model wins or ties.
In the first round, the best tested model scored about 47 percent, just below human level. A year earlier it had been around 14 percent. Typical mistakes were not a lack of knowledge, but sloppiness: wrong file format, ignored instructions, messy layout. It’s also important to note that these are clearly defined individual tasks. The test does not measure team coordination, follow-up questions, or responsibility for decisions.
GDPval in headlines and product promises
You’ll most often encounter GDPval in reports about new AI models. When a provider writes that their system is “nearly on par with experts” on professional tasks, this test is often behind it. The win rate also shows up as an argument in analyses by banks and consulting firms when it comes to productivity or personnel costs.
It makes sense to put such figures into context. GDPval tests tasks whose result can be delivered as a file. Craftwork, bedside care, or customer conversations fall outside its scope. A smaller portion of the test is publicly accessible, so researchers can verify it themselves. Anyone discussing AI and the labor market should therefore treat GDPval as one measuring instrument among several – useful, but not the final word.