METR

METR

METR is a US nonprofit organization that measurably tests how autonomously AI systems can carry out longer work tasks. It became known for the "task length" metric and for safety evaluations that major AI providers commission before releasing new models.

METR is a nonprofit research organization from the United States. It examines computer programs that can perform tasks autonomously and measures how well they succeed at this. The name stands for “Model Evaluation and Threat Research,” roughly: assessment of programs and research into possible dangers. The group emerged in 2023 from a team that had previously belonged to the research organization ARC, and has operated independently ever since. Providers such as OpenAI and Anthropic have given METR their new systems to test before release. METR is therefore neither a government agency nor a company, but an independent evaluator that those being tested bring in voluntarily.

The independent test bench before market launch

So far there are hardly any binding testing regulations for AI. For cars, lawmakers mandate crash tests; for medications, there are approval studies. For a new language model, it is usually the manufacturer itself that decides what to test and what to publish. This is exactly the gap METR steps into. When an external group examines a model, the result is more credible than a self-report.

The reports also matter for politics and regulatory bodies. Anyone wanting to write rules for AI needs figures on what the systems can actually do. Mere claims from marketing copy do not help with that. METR results are therefore cited in hearings, legislative debates, and in the providers' technical reports.

For investors and journalists, the organization is interesting for a different reason. Its measurements show how quickly capabilities change from year to year. This is one of the few reasonably sober foundations for assessing progress. However, METR remains dependent on companies voluntarily granting access. The organization has no right to conduct testing.

Task length as a benchmark

Classic tests quiz AI systems on knowledge, similar to an exam. METR takes a different approach and gives real work assignments. A system might be asked to find a software bug, analyze data, or get a piece of software running. In doing so, it is allowed to use tools, meaning it can open files, execute commands, and search the web. What is measured is whether the task is ultimately completed correctly.

The decisive trick is the conversion into time. METR first has experienced professionals solve the same tasks and records how long they take. It then checks which task length a model can still manage in roughly half of the cases. The result is a figure in minutes or hours: the so-called time-horizon metric. In effect, it indicates how long a human would have needed for the work that the model can just barely still handle.

This metric has become well known because it has risen sharply over the years. Where models used to manage only tasks lasting a few minutes, they can now handle noticeably longer pieces of work. One misunderstanding should be avoided here. The number does not describe how fast the AI computes, but how extensive the task is allowed to be. And it only applies to completed, clearly verifiable assignments, not to open-ended creative work.

METR in model reports and headlines

Most often, the name is encountered in the technical accompanying reports of new models. There it will state that an external group tested the system for risky capabilities before launch. Tests check, for example, whether a model can autonomously exploit security vulnerabilities or continue tasks for hours without human oversight. Such sections have appeared, among others, in models from OpenAI and Anthropic.

In the news, one usually encounters METR via two topics. The first are the task-length curves that appear in articles about the pace of AI development. The second is a 2025 study on the work of programmers. Experienced developers using AI assistance judged themselves to be faster, but when measured, actually took longer than without it. The result was widely discussed because it contradicts common product promises.

METR should not be confused with government bodies such as the AI safety institutes in the US or the UK. These are part of the administration, whereas METR is not. Its work also differs markedly from pure chatbot rankings, where often only which answer users prefer counts. METR, by contrast, is interested in whether a task was actually completed.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.