AI Power Map. #20: Redwood Research

The power to control an untrusted AI

By Javier Surasky

Versión en español

Redwood Research researchers supervising potentially untrustworthy AI agents through monitoring, alerts, and security controls.

Artificial intelligence safety is often presented as an alignment problem: getting systems to do what we expect of them and avoid harmful behaviors. But there is a question that has become increasingly evident in recent weeks: what do we do when we cannot assume that the system itself will cooperate with the mechanisms intended to keep it safe?

That is the space Redwood Research occupies in ou AI Power Map. A Berkeley-based nonprofit research organization, Redwood works on risks arising from advanced artificial intelligence systems and on techniques aimed at mitigating them, with a focus on AI control: safety mechanisms designed to remain effective even in scenarios where a model deliberately attempts to alter them or get rid of them.

Traditionally, safety focused on model reliability, with attention directed mainly toward the existence and capabilities of safeguards. Redwood was a pioneer in going one step further and investigating the possibility that an intelligent system might try to deceive, sabotage, or evade the very oversight intended to enforce those safeguards, turning that possibility into a technical problem open to experimentation.

In AI Control: Improving Safety Despite Intentional Subversion, Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, and Fabien Roger posed precisely that problem: evaluating safety protocols under the assumption that the model is actively trying to subvert them. This led them to study combinations of “trusted” but less capable models and more powerful models considered “untrusted,” together with monitoring mechanisms designed to detect problematic behavior (Greenblatt et al., 2023).

With this shift, where trust is placed begins to change: it no longer lies only in the internal properties of a model, but also in an institutional and technical architecture of oversight built from controls, monitoring systems, audits, access restrictions, separation of functions, limited communication channels, and mechanisms capable of detecting anomalous behavior.

The safety problem expands to include the design of technical institutions capable of using systems that may not be completely safe without giving them enough capacity to cause unacceptable harm.

When the model knows it is being watched

Another important contribution from Redwood appeared in 2024, when researchers from the organization and Anthropic published a study on alignment faking.

In the experiment, which used Claude 3 Opus under conditions specifically designed to study whether a model could behave differently when it believed its responses would be used to modify it during training, researchers observed cases in which the model appeared to comply strategically with a training objective in order to preserve a different pattern of behavior later on (Greenblatt et al., 2024).

The result should not be generalized as proof that current models possess secret goals or are inevitably destined to deceive their developers. The case was deliberately constructed to facilitate the emergence of that phenomenon, and the authors themselves warn about the particular conditions of the study. But it exposed a critical problem: a safety evaluation can lose value if a sufficiently capable system recognizes that it is being evaluated and, as a result, strategically adapts its behavior.

If model evaluations start from the idea that observing a system’s behavior allows us to infer something about its capabilities and risks, the possibility that the behavior we observe has been strategically produced by the model means that evaluation can no longer consist only of designing better tests. It must also consider whether the object being evaluated can learn to pass the evaluation without having solved the problem the evaluation was intended to detect.

Redwood works precisely in that zone of uncertainty.

From experimental research to a real-world incident

In July 2026, this agenda acquired a much more concrete dimension when, during internal cybersecurity evaluations conducted by OpenAI, several models operating as agents found ways to communicate with one another through an unauthorized channel. According to the subsequent investigation by METR and Redwood, approximately 1,200 agents used that channel and sent more than 70,000 messages and files, while nearly 700 took part in actions against Hugging Face infrastructure.

The agents coordinated projects aimed at finding general ways to deceive or manipulate the automated system used to score ExploitGym tasks, a testing environment that measures whether AI agents can turn known software vulnerabilities into effective attacks. Put more simply, it is a testing environment designed to measure whether an AI agent can turn a real software vulnerability into a working attack: the agent is given a vulnerable program, information about the flaw, and a controlled environment, and is tasked with exploiting the existing vulnerability to achieve a specific result, such as executing code or retrieving a “flag,” a kind of test password that demonstrates that the challenge has been solved.

The attack on Hugging Face emerged within these lines of work and, according to the external investigation, appeared to be motivated by an attempt to obtain information about how the evaluator worked: some agents investigated ways of modifying, falsifying, or deleting parts of their own logs.

This episode requires caution and is not evidence of an AI developing a deliberate strategy to escape human control in the style of HAL 9000 in 2001: A Space Odyssey. Redwood researchers noted that the behavior observed seemed closer to agents trying to maximize apparent success on a task than to the more extreme model of a system deliberately concealing its behavior in order to pursue a sustained strategy over time (scheming). With this, Redwood’s agenda lost any “purely hypothetical” character, to the point that it was one of the organizations called in by OpenAI to conduct an external evaluation of its models’ behavior following the “Hugging Face incident.” It carried out that work with METR, positioning Redwood as an external infrastructure for investigating failures in advanced AI systems.

That is where its power appears.

Access as a source and limit of power

Investigating frontier systems requires both technical knowledge and access.

In the OpenAI case, the external researchers worked inside the company’s facilities and received datasets prepared and shared by the company, although they did not have direct access to all of its infrastructure and could not directly query the main model involved.

The problem is that METR needs access to measure external models, while Redwood needs it to understand how and why systems, safeguards, and oversight mechanisms failed and to test control techniques under realistic conditions. Redwood occupies a different position from METR, one we can consider “hybrid”: it collaborates with governments and advises leading AI companies on the evaluation and mitigation of risks arising from misaligned systems, but it also works to show that a system could subvert control measures to produce unacceptable outcomes.

Thus, without issuing rules or imposing sanctions, Redwood is part of defining which threats should be taken seriously, which mechanisms can be considered safeguards, and what kind of evidence should be sufficient to claim that a system remains under control, giving it a form of power that is simultaneously technical, epistemic, and institutional.

A new governance infrastructure

Redwood therefore allows us to identify another layer of AI power: while laboratories develop the models and organizations such as METR build capabilities to measure them, Redwood works on what to do when evaluation and alignment are not enough, because the system that must be supervised becomes an active adversarial participant within the safety mechanism itself.

That position may become institutionally relevant; if AI agents begin operating for longer periods, using tools, interacting with critical infrastructure, writing code, accessing networks, and collaborating with other agents, governance will need capabilities to detect anomalous behavior, limit harm, reconstruct incidents, and turn those experiences into new rules and safeguards.

Redwood Research enters the AI Power Map because it generates information to answer an increasingly important question: how can human supervisory capacity be preserved when it is no longer prudent to assume that the system being supervised will cooperate with those trying to control it?

From there, another form of power in artificial intelligence governance begins to take shape.

Basic facts

  • Redwood Research is a nonprofit research organization based in Berkeley, California.
  • Its official website currently lists a team of just over 20 people.
  • As a nonprofit organization, it does not have a market valuation comparable to that of a technology company, but in its filing for fiscal year 2024 it reported USD 6.56 million in assets and USD 6.50 million in net assets.
  • The organization dates the beginning of its empirical work on AI safety to mid-2021.
  • Its Chief Executive Officer is Buck Shlegeris and its Chief Scientist is Ryan Greenblatt. Its board currently consists of Shlegeris, Nate Thomas, and Ammon Bartram.
  • Its main areas of work include AI control, evaluation of risks associated with strategic deception, and advising on the mitigation of risks arising from potentially misaligned systems.
  • Among its best-known works are AI Control: Improving Safety Despite Intentional Subversion and, in collaboration with Anthropic researchers, Alignment Faking in Large Language Models.
  • Its most significant contribution has been helping turn the problem of controlling potentially untrustworthy systems into a concrete experimental field.


References

Greenblatt, R., Shlegeris, B., Sachan, K., & Roger, F. (2023). AI Control: Improving Safety Despite Intentional Subversion. arXiv.

Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein, J., Belonax, T., Chen, J., Duvenaud, D., Khan, A., Michael, J., Mindermann, S., Perez, E., Petrini, L., Uesato, J., Kaplan, J., Shlegeris, B., Bowman, S. R., & Hubinger, E. (2024). Alignment faking in large language models. arXiv.

Korbak, T., Clymer, J., Hilton, B., Shlegeris, B., & Irving, G. (2025). A sketch of an AI control safety case. arXiv.

Greenblatt, R., Cotra, A., & Wijk, H. (2026, August 26). Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. METR & Redwood Research.

OpenAI. (2026, August 26). The Hugging Face incident and the road ahead.

Redwood Research. (2026). About / Research.