Proposal for tracking the effects of architecture on monitorability
How AI companies could be transparent about monitorability-relevant evidence and policies.
Research notes, results, and thinking from the Redwood team — collected here from our Substack, LessWrong, and the Alignment Forum so you can read it all in one place.
How AI companies could be transparent about monitorability-relevant evidence and policies.

"Serial depth between text bottlenecks" as a proxy for latent reasoning abilities.

We recently published the report from our brief independent investigation into this incident. You can read the full report here.

Unsanctioned coordination, like we saw in the Hugging Face incident, could enable future AIs to take over

Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned, it…

When a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult

We need more details

And what the incident can’t tell us about alignment

What are the broader lessons from this incident?

Yes, but less than had the models been schemers.

We recently ran a strategy fellowship through Astra. As part of this, we ran a reading group for our fellows on some of the topics that we think are important for thinking about AI futurism (key dynamics in AI development, existential risk…

If it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement model.

Models' no-CoT time horizon has doubled roughly every year.

When is "increasing safety budget" a useful concept?

We’ve just released a new paper: Retrying vs Resampling in AI Control. We revisit the resampling protocols introduced in Ctrl-Z with an up-to-date setting and much stronger models, and compare them against “retrying” protocols similar to…

We’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are…

Full automation likely yields a one-time speed-up and higher returns from compute

Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen:

Deployment-time spread is the most plausible near-term route to consistent adversarial misalignment

My median guess: it's as good as a crystal ball that sees 2.5 months into the future.

Last week, OpenAI staff shared an early draft of Investigating the consequences of accidentally grading CoT during RL with Redwood Research staff.
Fitness-seeking is increasingly what misalignment looks like in practice—how should we respond?

One of the main hopes for AI safety is using AIs to automate AI safety research. However, if models are misaligned, then they may sabotage the safety research. For example, misaligned AIs may try to:

Eliciting long-term forecasts from myopic fitness-seekers

A controlled reward-seeking motivation could make AI safer and more useful

Third-party experts should assess defenses against tampering and theft — and publish high-level findings

We might want to strike deals with early misaligned AIs in order to reduce takeover risk and increase our chances of reaching a better future. For example, we could ask a schemer who has been undeployed to review its past actions and point…

A new control setting for more realistic software engineering deployments
Authors: Dylan Xu, Alek Westover, Vivek Hebbar, Sebastian Prasanna, Nathan Sheffield, Buck Shlegeris, Julian Stastny

In my experience, AIs often oversell their work, downplay problems, and cheat

Safely navigating the intelligence explosion will require much more careful development

We study trusted monitoring for AI control, where a weaker trusted model reviews the actions of a stronger untrusted agent and flags suspicious behavior for human audit. We propose a simple mathematical model relating safety (true positive…

Better estimates of uplift at AI companies seem helpful
Thanks to Buck Shlegeris for feedback on a draft of this post.

My predictions about what is going on right now

I've updated towards substantially shorter timelines

A common element in many AI control schemes is monitoring – using some model to review actions taken by an untrusted model in order to catch dangerous actions if they occur. Monitoring can serve two different goals. The first is detection:…
They'd renege on non-binding commitments, defect against copies of themselves in prisoner's dilemmas, etc.

AI inference is still cheap relative to human labor

RL would encourage on-episode reward seeking, but beyond-episode reward seekers may learn to goal-guard.

Some unintended preferences are cheap to satisfy, and failing to satisfy them needlessly turns a cooperative situation into an adversarial one.

The executive branch can, and probably would, block a frontier AI company's departure.

Berkeley, April 18-19
Current LLMs externalize lots of their reasoning in human interpretable language. This reasoning is sometimes unfaithful, sometimes strange and concerning, and LLMs can do somewhat impressive reasoning without using CoT, but my overall…

Reward-seekers are supposed to be safer because they respond to incentives under developer control. But what if they also respond to incentives that aren't?

How can we make AIs aligned and well-elicited on extremely hard to check open ended tasks?

Most recent progress probably isn't from unsustainable inference scaling
Some people hope we can force misaligned AI systems to automate alignment research by controlling them. I'll call this the "diffuse control" plan. This plan has several unfortunate properties:

If you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren’t the same.
I'm worried about AI models intentionally doing bad things, like sandbagging when doing safety research. In the regime where the AI has to do many of these bad actions in order to cause an unacceptable outcome, we have some hope of…

With Buck Shlegeris and Ryan Greenblatt
In order to control an AI model's worst-case performance, we need to understand its generalization properties to situations where it hasn't been trained. It seems plausible that powerful AI models will Fake Alignment and then generalize…

Recent AIs are much better at chaining together knowledge in a single forward pass

Opus 4.5 has around a 3.5 minute 50%-reliablity time horizon
AI control tries to ensure that malign AI models can’t cause unacceptable outcomes even if they optimize for such outcomes. AI control evaluations use a red-team–blue-team methodology to measure the efficacy of a set of control measures.…

AI can sometimes distribute cognition over many extra tokens

We’ve just released BashArena, a new high-stakes control setting we think is a major improvement over the settings we’ve used in the past. In this post we’ll discuss the strengths and weaknesses of BashArena, and what we’ve learned about…

The basic arguments about AI motivations in one causal graph

A reason alignment could be hard

I operationalize Anthropic's prediction of "powerful AI" and explain why I'm skeptical

And this seems caused by training on alignment evals.

There is some concern that training AI systems on content predicting AI misalignment will hyperstition AI systems into misalignment. This has been discussed previously by a lot of people: Anna Salamon, Alex Turner, the AI Futures Project,…

I'm skeptical that Dario's prediction of AIs writing 90% of code in 3-6 months has come true

Can we study scheming by studying AIs trained to act like schemers?

A strategy for handling scheming

It's a promising design for reducing model access inside AI companies.

Different plans for different levels of political will

Large fractions of people will die, but literal human extinction seems unlikely

Transparency about just safety cases would have bad epistemic effects

Studying actual schemers seems promising but tricky

An initial narrow proposal

AIs that speed up engineering by 2x wouldn't accelerate AI progress that much

Above trend progress due to a rapid increase in RL env quality is unlikely

System cards are established but other approaches seem importantly better

More thoughts on making deals with schemers

When and why we should refrain from lying

AGI before 2029 now seems substantially less likely

To wit: LLM APIs, agent scaffolds, code review, and detection-and-response systems

Maybe we should (re)assess the case for relatively fast progress after the GPT-5 release.

It's like making challenging evals, but more constrained

Empirical AI security/safety projects across a variety of areas
(Last updated Jul 28th 2025)

And what if you have both at once?
Ryan’s podcast with Rob Wiblin has just come out! I think it turned out great. I particularly enjoyed Ryan’s discussion of different pathways to AI takeover, which I don’t think has been discussed in as much depth elsewhere. I really…

And what does this mean for AI control?

We should study methods to train away deeply ingrained behaviors in LLMs that are structurally similar to scheming.

"They need to act aligned" often isn't precise enough

How much time can control buy us during the intelligence explosion?

Once AIs match top humans, what are the returns to further scaling and algorithmic improvement?

And why I think insider threat from AI combines the hard parts of both problems.

...could help us to prevent takeover attempts from more dangerous misaligned AIs created later.

Training the policy to not do egregious bad actions we detect has downsides and we might be able to do better

Distillation is cheap; how can we use it to improve safety?

Can a scheming AI's goals really stay unchanged through training?

Defending against alignment problems that might come with long-term memory

Some reasons why relatively weak AIs might still be important when we have very powerful AIs

A new analysis of the risk of AIs intentionally performing poorly.

Clarifying ways in which faking alignment during training is neither necessary nor sufficient for the kind of scheming that AI control tries to defend against.

My views on what's driving AI progress and where it's headed.

Preventing research sabotage will require techniques very different from the original control paper.

A list of easy-to-start directions in AI control targeted at independent researchers without as much context or compute

(There are a few)

A model of the relationship between higher level goals, explicit reasoning, and learned heuristics in capable agents.

What if getting strong evidence of scheming isn't the end of your scheming problems, but merely the middle?

A new paper on AI control for agents.

Evidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing.

A breakdown of how higher capabilities increase risk

What are all the research and implementation areas helpful for control?

What methods can we use to ensure control?
My podcast with Rob Wiblin from 80,000 Hours just came out. I’m really happy with how it turned out. I talked about a bunch of stuff on the podcast that I don’t think we’ve written up before.

How can we prevent AIs from intentionally underperforming on our metrics?

What are the methods and issues when failures occur diffusely over many actions?

What are the main threats and how should we prioritize them?

Developing AI employees that are safer than human ones

Insights from a long technical paper compressed into a fun little commentary

Are we ready for this?

A scary scenario that's worth planning for

For a summary of this post see the thread on X.

A quantitative description of how I expect to change my mind.

Why are we so friendly to the red team?

Buck Shlegeris and Ryan Greenblatt originally motivated control evaluations as a way to mitigate risks from ‘scheming’ AI models: models that consistently pursue power-seeking goals in a covert way; however, many adversarial model…

The complement to control evaluations

In our experiments, AIs will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences.

There are crucial disanalogies between preventing jailbreaks and preventing misalignment-induced catastrophes.

In this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective.

One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT). The basic version of this strategy is something like: Before…

And diagrams describing how threat scenarios involving misaligned AI involve compromising the system in different places.

Suppose you’ve trained a really clever AI model, and you’re planning to deploy it in an agent scaffold that allows it to run code or take other actions. You’re worried that this model is scheming, and you’re worried that it might only need…

I'm not so sure.

Is AI takeover like a nuclear meltdown? A coup? A plane crash?

You can just draw more samples

AI might be really helpful for reducing security risk.

It’s interesting to classify possible AI catastrophes based on whether or not they involve a "rogue deployment".

Unlike most files you might want to secure, model weights are extremely big. This might make them much easier to secure.

The easiest way to be sure an AI isn't scheming against you is to note that it's too dumb to pull that off. What happens if that's the only way we have to rule out scheming?

How could an AI lab serving AIs to customers manage catastrophic misuse without solving adversarial robustness?

If your AIs are trying to escape, it's crucial to think about whether you can catch them before they succeed, because catching them red-handed gives you lots of options you didn't have before.

Labs should make sure that powerful models can't cause unacceptably bad outcomes even if the AIs try to.