By Javier Surasky
The new power of measuring AI safety
In the
governance of artificial intelligence, power is expressed through laws,
treaties, and corporate decisions, but another layer operates silently, hidden
from non-expert users: the capacity to measure.
Whoever defines what is evaluated, how it is evaluated and what counts as risk gains influence over the future of this technology. That is the space in which METR
operates. METR stands for Model Evaluation and Threat Research, an
organization that seeks to turn the question of when an AI system begins to
pose a serious social risk into a technical program.
METR
presents itself as a nonprofit organization dedicated to scientifically
measuring whether AI systems may threaten society with catastrophic harms, and
to developing methods that make it possible to make better decisions about
their development (METR, 2026a). Its focus is on analyzing whether a model can
perform substantive tasks autonomously: conducting research, programming,
operating in digital environments, solving long-running problems, or developing
capabilities that increase safety risks.
METR is not
a formal regulator that issues rules or imposes sanctions. Its power is
epistemic, technical, and institutional, and rests on producing evidence
about emerging AI capabilities: companies, governments, safety institutes, regulators, and the media can use that evidence. In that way, METR
influences the criteria by which others decide how to govern AI because, as we
already know, in AI, measuring is not a neutral task: it means prioritizing
certain harms, certain scenarios and certain forms of intervention.
One of
METR’s major contributions is to move the conversation away from abstract
debates about “existential risk” and toward concrete tests of autonomy, agency, and performance in long-horizon tasks, using evaluation resources that include
task sets, software tools, and methodological guidance for measuring autonomous
system capabilities (METR, 2026b). In other words, it asks whether a model can
act effectively over extended periods, chain steps together, use tools, and
produce results comparable to those of a trained person.
The HCAST
benchmark illustrates this kind of reasoning well: instead of measuring
isolated answers, HCAST proposes human-calibrated software tasks to compare the
performance of AI agents against tasks that trained people complete within
estimated timeframes, in order to know whether an agent can be trusted to
complete a task that would take a human a given amount of time (Rein et al.,
2025).
In the AI
ecosystem, METR thus occupies an intermediate space between technical civil
society, independent auditing, and anticipatory governance, with an authority
that depends on scientific credibility.
To fulfill
its mission, METR requires access to frontier models. It needs companies to let it test powerful systems, sometimes before release or with information not available to the public. This creates an institutional tension: the
more it depends on the companies being evaluated for access to models,
information or infrastructure, the more important it becomes to demonstrate
independence. Without access, external evaluation can become superficial, but
conditioned access risks soft capture by collaborating companies.
METR seems
aware of this problem when it raises the need for clearer standards, better
ways to verify information reported by companies, and more frequent periodic
evaluations (METR, 2026c).
For States,
METR represents a potential source of technical capacity, providing
infrastructure, specialized talent, and access to advanced models they lack.
In this
context, METR also ends up operating as an intermediary between the technical
world and political decision-makers. But, once again, we must introduce a warning from a democratic standpoint: AI governance cannot be reduced to a conversation among private labs, technical experts, and specialized evaluators, with political decision-makers depending on their conclusions because they lack the capacity to understand models and access data.
Moreover,
catastrophic risks do not exhaust the AI safety agenda, because current
harms such as discrimination, market concentration, surveillance, precarious
work, technological dependence, environmental impacts and inequalities in
access to computational infrastructure also matter, and they are not the core focus
of METR’s work.
For all
these reasons, METR should be read as an emerging power actor that helps define
the categories through which AI will be controlled, without claiming to control
it.
In the AI
ecosystem, measuring is a form of intervention, and that is where METR is
situated as a reference actor.
Basic facts
- METR (Model Evaluation and Threat Research) is based in California.
- It was created in December 2023, when ARC Evals announced that it was ending its incubation period within the Alignment Research Center and adopting the name METR as an independent nonprofit organization (METR, 2023).
- Its founder and executive director is Beth Barnes, who previously worked at DeepMind and OpenAI.
- It is funded mainly through donations. It has declared support from The Audacious Project, individuals, the Sijbrandij Foundation, Pew Charitable Trusts, Schmidt Sciences, the Packard Foundation, the LaCentra-Sumerlin Foundation, the Astralis Foundation, Expa.org, the AI Security Institute, Longview Philanthropy, Effektiv Spenden, the Survival and Flourishing Fund and individual donors. A smaller share of its income comes from a technical assistance contract with the European AI Office, and it does not accept funding from frontier AI companies or their employees, although it acknowledges that those companies provide it with free tokens for evaluations, research and engineering.
- Its strongest contribution lies in showing that AI safety requires institutions capable of evaluating real capabilities, not corporate promises or public impressions.
- Its limit lies in the fact that it evaluates, but does not seek to replace political decisions about which risks are acceptable, who participates in that definition and what consequences should be triggered when a company crosses certain thresholds.
References
METR.
(2023, December 4). ARC Evals is now METR. https://metr.org/blog/2023-12-04-metr-announcement/
METR.
(2026a). About METR. https://metr.org/about
METR.
(2026b). Resources for measuring autonomous AI capabilities. https://metr.org/measuring-autonomous-ai-capabilities/
METR.
(2026c, May 19). Frontier Risk Report (February to March 2026). https://metr.org/blog/2026-05-19-frontier-risk-report/
Rein, D., Becker, J., Deng, A., Nix, S., Canal, C.,
O’Connell, D., Arnott, P., Bloom, R., Broadley, T., Garcia, K., Goodrich, B.,
Hasin, M., Jawhar, S., Kinniment, M., Kwa, T., Lajko, A., Rush, N., Sato, L. J.
K., Von Arx, S., Gleave, A., & Barnes, E. (2025). HCAST: Human-Calibrated Autonomy Software Tasks. METR. https://metr.org/hcast.pdf
.png)