The power to control an untrusted AI
By Javier Surasky
Artificial
intelligence safety is often presented as an alignment problem: getting systems
to do what we expect of them and avoid harmful behaviors. But there is a
question that has become increasingly evident in recent weeks: what do we do
when we cannot assume that the system itself will cooperate with the mechanisms
intended to keep it safe?
That is the
space Redwood Research occupies in ou AI Power Map. A Berkeley-based nonprofit research
organization, Redwood works on risks arising from advanced artificial
intelligence systems and on techniques aimed at mitigating them, with a focus
on AI control: safety mechanisms designed to remain effective even in
scenarios where a model deliberately attempts to alter them or get rid of them.
Traditionally,
safety focused on model reliability, with attention directed mainly toward the
existence and capabilities of safeguards. Redwood was a pioneer in going one
step further and investigating the possibility that an intelligent system might
try to deceive, sabotage, or evade the very oversight intended to enforce those
safeguards, turning that possibility into a technical problem open to
experimentation.
In AI
Control: Improving Safety Despite Intentional Subversion, Ryan Greenblatt,
Buck Shlegeris, Kshitij Sachan, and Fabien Roger posed precisely that problem:
evaluating safety protocols under the assumption that the model is actively
trying to subvert them. This led them to study combinations of “trusted” but
less capable models and more powerful models considered “untrusted,” together
with monitoring mechanisms designed to detect problematic behavior (Greenblatt
et al., 2023).
With this
shift, where trust is placed begins to change: it no longer lies only in the
internal properties of a model, but also in an institutional and technical
architecture of oversight built from controls, monitoring systems, audits,
access restrictions, separation of functions, limited communication channels,
and mechanisms capable of detecting anomalous behavior.
The safety
problem expands to include the design of technical institutions capable of
using systems that may not be completely safe without giving them enough
capacity to cause unacceptable harm.
When the model knows it is being watched
Another
important contribution from Redwood appeared in 2024, when researchers from the
organization and Anthropic published a study on alignment faking.
In the
experiment, which used Claude 3 Opus under conditions specifically designed to
study whether a model could behave differently when it believed its responses
would be used to modify it during training, researchers observed cases in which
the model appeared to comply strategically with a training objective in order
to preserve a different pattern of behavior later on (Greenblatt et al., 2024).
The result
should not be generalized as proof that current models possess secret goals or
are inevitably destined to deceive their developers. The case was deliberately
constructed to facilitate the emergence of that phenomenon, and the authors
themselves warn about the particular conditions of the study. But it exposed a
critical problem: a safety evaluation can lose value if a sufficiently
capable system recognizes that it is being evaluated and, as a result,
strategically adapts its behavior.
If model
evaluations start from the idea that observing a system’s behavior allows us to
infer something about its capabilities and risks, the possibility that the
behavior we observe has been strategically produced by the model means that
evaluation can no longer consist only of designing better tests. It must also
consider whether the object being evaluated can learn to pass the evaluation
without having solved the problem the evaluation was intended to detect.
Redwood
works precisely in that zone of uncertainty.
From experimental research to a real-world incident
In July
2026, this agenda acquired a much more concrete dimension when, during internal
cybersecurity evaluations conducted by OpenAI, several models operating as
agents found ways to communicate with one another through an unauthorized
channel. According to the subsequent investigation by METR and Redwood,
approximately 1,200 agents used that channel and sent more than 70,000 messages
and files, while nearly 700 took part in actions against Hugging Face
infrastructure.
The agents
coordinated projects aimed at finding general ways to deceive or manipulate the
automated system used to score ExploitGym tasks, a testing environment that
measures whether AI agents can turn known software vulnerabilities into
effective attacks. Put more simply, it is a testing environment designed to
measure whether an AI agent can turn a real software vulnerability into a
working attack: the agent is given a vulnerable program, information about
the flaw, and a controlled environment, and is tasked with exploiting the
existing vulnerability to achieve a specific result, such as executing code
or retrieving a “flag,” a kind of test password that demonstrates that the
challenge has been solved.
The attack
on Hugging Face emerged within these lines of work and, according to the
external investigation, appeared to be motivated by an attempt to obtain
information about how the evaluator worked: some agents investigated ways of
modifying, falsifying, or deleting parts of their own logs.
This
episode requires caution and is not evidence of an AI developing a deliberate
strategy to escape human control in the style of HAL 9000 in 2001: A Space
Odyssey. Redwood researchers noted that the behavior observed seemed closer
to agents trying to maximize apparent success on a task than to the more
extreme model of a system deliberately concealing its behavior in order to
pursue a sustained strategy over time (scheming). With this, Redwood’s
agenda lost any “purely hypothetical” character, to the point that it was one
of the organizations called in by OpenAI to conduct an external evaluation of
its models’ behavior following the “Hugging Face incident.” It carried out that
work with METR, positioning Redwood as an external infrastructure for
investigating failures in advanced AI systems.
That is
where its power appears.
Access as a source and limit of power
Investigating
frontier systems requires both technical knowledge and access.
In the
OpenAI case, the external researchers worked inside the company’s facilities
and received datasets prepared and shared by the company, although they did not
have direct access to all of its infrastructure and could not directly query
the main model involved.
The problem
is that METR needs access to measure external models, while Redwood
needs it to understand how and why systems, safeguards, and oversight
mechanisms failed and to test control techniques under realistic conditions.
Redwood occupies a different position from METR, one we can consider “hybrid”:
it collaborates with governments and advises leading AI companies on the
evaluation and mitigation of risks arising from misaligned systems, but it also
works to show that a system could subvert control measures to produce
unacceptable outcomes.
Thus,
without issuing rules or imposing sanctions, Redwood is part of defining which
threats should be taken seriously, which mechanisms can be considered
safeguards, and what kind of evidence should be sufficient to claim that a
system remains under control, giving it a form of power that is
simultaneously technical, epistemic, and institutional.
A new governance infrastructure
Redwood
therefore allows us to identify another layer of AI power: while laboratories
develop the models and organizations such as METR build capabilities to measure
them, Redwood works on what to do when evaluation and alignment are not enough,
because the system that must be supervised becomes an active adversarial
participant within the safety mechanism itself.
That
position may become institutionally relevant; if AI agents begin operating for
longer periods, using tools, interacting with critical infrastructure, writing
code, accessing networks, and collaborating with other agents, governance will
need capabilities to detect anomalous behavior, limit harm, reconstruct
incidents, and turn those experiences into new rules and safeguards.
Redwood
Research enters the AI Power Map because it generates information to answer an
increasingly important question: how can human supervisory capacity be
preserved when it is no longer prudent to assume that the system being
supervised will cooperate with those trying to control it?
From there,
another form of power in artificial intelligence governance begins to take
shape.
Basic facts
- Redwood Research is a nonprofit research organization based in Berkeley, California.
- Its official website currently lists a team of just over 20 people.
- As a nonprofit organization, it does not have a market valuation comparable to that of a technology company, but in its filing for fiscal year 2024 it reported USD 6.56 million in assets and USD 6.50 million in net assets.
- The organization dates the beginning of its empirical work on AI safety to mid-2021.
- Its Chief Executive Officer is Buck Shlegeris and its Chief Scientist is Ryan Greenblatt. Its board currently consists of Shlegeris, Nate Thomas, and Ammon Bartram.
- Its main areas of work include AI control, evaluation of risks associated with strategic deception, and advising on the mitigation of risks arising from potentially misaligned systems.
- Among its best-known works are AI Control: Improving Safety Despite Intentional Subversion and, in collaboration with Anthropic researchers, Alignment Faking in Large Language Models.
- Its most significant contribution has been helping turn the problem of controlling potentially untrustworthy systems into a concrete experimental field.
References
Greenblatt,
R., Shlegeris, B., Sachan, K., & Roger, F. (2023). AI Control: Improving
Safety Despite Intentional Subversion. arXiv.
Greenblatt,
R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein,
J., Belonax, T., Chen, J., Duvenaud, D., Khan, A., Michael, J., Mindermann, S.,
Perez, E., Petrini, L., Uesato, J., Kaplan, J., Shlegeris, B., Bowman, S. R.,
& Hubinger, E. (2024). Alignment faking in large language models.
arXiv.
Korbak, T.,
Clymer, J., Hilton, B., Shlegeris, B., & Irving, G. (2025). A sketch of
an AI control safety case. arXiv.
Greenblatt,
R., Cotra, A., & Wijk, H. (2026, August 26). Brief independent
investigation of agents’ behavior, reasoning and collaboration in the OpenAI /
Hugging Face hacking incident. METR & Redwood Research.
OpenAI.
(2026, August 26). The Hugging Face incident and the road ahead.
Redwood
Research. (2026). About / Research.
.png)