
GDPval-AA v2
GDPval-AA v2 is the second version of a test collection used to assess how well AI systems perform real professional tasks. Instead of multiple-choice questions, the system receives work assignments, and experts then evaluate the result.
GDPval-AA v2 is a test for computer programs that process language and solve tasks. Such programs are called AI models. The test does not consist of knowledge questions, but of real work assignments from professional life: analyzing a spreadsheet, reviewing a draft contract, building a presentation. The model delivers a result, and people with professional experience judge whether the result is usable. The suffix “AA” stands for this procedure with human evaluation, the “v2” for the second, revised version of the task collection. The underlying idea comes from OpenAI’s GDPval approach, which collects tasks from many economic sectors.
Why professional tasks are more meaningful than exam questions
For a long time, AI models were evaluated primarily using exam questions. A model solves math problems or answers questions from a medical exam. Such tests have a drawback: the correct answers are often available somewhere on the internet. A model that has seen these texts during training can perform well without truly having the underlying ability. Experts call this data contamination.
Professional tasks are harder to game. There is rarely just one correct solution. A financial report can be calculated cleanly and yet be written incomprehensibly. Exactly such differences become apparent when a human from the relevant field checks the result. This means the test measures something that actually matters to companies.
For investors and editorial teams, this is the real point. The question is not whether a model passes an exam. The question is what share of paid work can be accomplished with it. Results from tests like GDPval-AA v2 are therefore frequently cited in discussions about the labor market. However, they should be read as a snapshot, not as a forecast.
From work assignment to evaluation score
The process has three stages. First, experts collect tasks from their everyday work, each with the necessary documents. Then several AI models work on the same tasks under identical conditions. Finally, reviewers compare the results, often without knowing which model produced what. This procedure is called blind evaluation and prevents a well-known brand name from coloring the judgment.
Pairwise comparison is frequently used. The reviewer sees two solutions side by side and decides which one is better. Many such comparisons produce a ranking. The best-known metric is the win rate: the proportion of cases in which the model’s solution is preferred over the comparison solution. A win rate of 50 percent means parity with the benchmark.
The second version was revised mainly to address the weaknesses of the first. Typical changes include clearer task descriptions, more consistent evaluation rules, and additional tasks from underrepresented industries. This has an inconvenient consequence: values from v1 and v2 cannot be directly compared. Anyone who places both versions side by side in a chart easily creates a false impression of progress.
Where the numbers show up
GDPval-AA v2 is most commonly encountered in announcements of new models. Vendors show bar charts in which their new system outperforms both their own predecessor model and the competition. Trade media pick up these numbers, sometimes without mentioning the test conditions. A critical look is always worthwhile: how many tasks, which industries, who did the evaluating?
The test also plays a role within companies. Anyone considering deploying AI in accounting or customer service is looking for reliable indicators. Results from job-related tests are more useful for this than scores from school exams. Nevertheless, they cannot replace a company’s own practical test, because every company has different documents and workflows.
A common misconception is that a high win rate means the professions in question are obsolete. The test measures individual, clearly defined tasks. A profession, however, consists of coordination, responsibility, and decisions made under uncertainty. That is exactly what GDPval-AA v2 does not capture.