CORE-Bench

CORE-Bench

CORE-Bench is a benchmark — that is, a standardized test — that measures how well AI systems can independently retrace and reproduce scientific research findings. It thus assesses a capability that is particularly critical for the use of AI in science.

Scientific studies are only reliable if others can repeat them and arrive at the same result. This is called reproducibility. CORE-Bench is a test that measures whether AI systems are capable of this: they are supposed to take the code and data of a real study and independently recompute the published results. The benchmark dates from 2024 and comprises hundreds of tasks from various scientific disciplines, including computer science, economics, and medicine. Each task is based on a real, published study with publicly available code.

Reproducibility as a yardstick for scientific AI

Many AI systems can summarize texts or answer questions. It’s a different matter entirely to trace the computational path of a study without gaps. To do this, the system must execute code, install dependencies, fix errors, and finally check whether the numbers are correct. This is a combination of technical understanding and planning-oriented action.

Why this matters: Many studies are already hard for humans to reproduce — missing packages, undocumented steps, outdated software. If an AI system could reliably manage this, it could take a considerable amount of this work off researchers' hands. CORE-Bench makes this progress measurable and comparable for the first time.

Structure and difficulty levels of the test

CORE-Bench divides its tasks into three difficulty levels. The easiest level only asks whether the system understands which results a study actually contains. The middle level requires the system to execute the code, but the environment — meaning programs and libraries — is already set up. At the hardest level, the system must set up everything from scratch itself and then deliver the correct figures.

What is evaluated is the result, not the path taken. If the values calculated by the AI system match the values from the original study, the task is considered solved. That sounds simple, but it isn’t: even the most capable systems fail on a large portion of the tasks at the highest level. When the benchmark was released, the best system tested solved about 21 percent of the hardest tasks — a sobering result that shows how far AI still is from genuine scientific autonomy.

The benchmark employs so-called AI agents — that is, systems that don’t just output text but actively use tools, write and execute code, and make decisions independently. Classic chatbots would fundamentally be unsuited to this task.

CORE-Bench in the research debate and in products

CORE-Bench comes up mainly in discussions about so-called AI agents — systems designed to independently complete multi-step tasks. Companies like OpenAI, Google, and Anthropic are developing such agents and need objective tests to document their progress. CORE-Bench is one of the few benchmarks that poses a clearly verifiable task from the real world: either the numbers are correct, or they aren’t.

In science policy, CORE-Bench is used as an argument in the debate over the so-called replication crisis. Many fields, especially psychology and parts of medicine, struggle with the fact that published studies often cannot be confirmed. If AI systems could automatically check reproducibility, this would be a concrete tool against this problem.

For readers of tech news, CORE-Bench is usually encountered when new agent models are introduced and manufacturers report their results on various benchmarks. It is considered particularly meaningful because it does not pose multiple-choice questions but demands real, complex work.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.