Blog

Research notes, results, and thinking from the Redwood team — collected here from our Substack, LessWrong, and the Alignment Forum so you can read it all in one place.

Filter by topic
Visit our Substack
138 posts

Proposal for tracking the effects of architecture on monitorability

How AI companies could be transparent about monitorability-relevant evidence and policies.

Risk assessmentControl & Monitoring
Redwood ResearchSeptember 10, 2026

An operationalization of opaque serial depth

"Serial depth between text bottlenecks" as a proxy for latent reasoning abilities.

Risk assessmentControl & Monitoring
Nathan Sheffield, Alek Westover, Lukas Finnveden, Alexa Pan, Julian Stastny, Ryan GreenblattSeptember 10, 2026

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident

We recently published the report from our brief independent investigation into this incident. You can read the full report here.

Risk assessment
Ryan GreenblattAugust 27, 2026

AI swarms are starting to pose indirect takeover risk

Unsanctioned coordination, like we saw in the Hugging Face incident, could enable future AIs to take over

Threat modeling
Oak Hu, Alex MallenAugust 12, 2026

SOTA alignment assessments don’t strongly update us against misalignment

Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned, it…

Risk assessment
Alexa PanJuly 31, 2026

Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs

When a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult

Control & Monitoring
Caleb Biddulph, Adam KaufmanJuly 27, 2026

An OpenAI model left notes about how to evade containment

We need more details

Risk assessment
Alex MallenJuly 26, 2026

The OpenAI models that hacked Hugging Face weren’t just following instructions

And what the incident can’t tell us about alignment

Risk assessment
Girish GuptaJuly 25, 2026

The OpenAI/Huggingface incident | Redwood Research podcast episode 2

What are the broader lessons from this incident?

Risk assessment
Ryan Greenblatt, Buck ShlegerisJuly 23, 2026

Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?

Yes, but less than had the models been schemers.

Risk assessmentThreat modeling
Alex Mallen, Girish GuptaJuly 23, 2026

AI Futurism Reading List

We recently ran a strategy fellowship through Astra. As part of this, we ran a reading group for our fellows on some of the topics that we think are important for thinking about AI futurism (key dynamics in AI development, existential risk…

AI strategy and futurism
Alexa PanJuly 2, 2026

The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn't

If it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement model.

Science & Mitigations
Alek Westover, Alexa Pan, Sebastian Prasanna, Arun JoseJune 18, 2026

Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models

Models' no-CoT time horizon has doubled roughly every year.

Science & Mitigations
Anders Cairns WoodruffJune 10, 2026

Efficient tradeoffs and the safety-usefulness tradeoff model

When is "increasing safety budget" a useful concept?

AI strategy and futurism
Buck ShlegerisJune 8, 2026

Retrying vs Resampling in AI Control

We’ve just released a new paper: Retrying vs Resampling in AI Control. We revisit the resampling protocols introduced in Ctrl-Z with an up-to-date setting and much stronger models, and compare them against “retrying” protocols similar to…

Control & Monitoring
James Lucassen, Adam KaufmanMay 29, 2026

Advice for making robust-to-training model organisms

We’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are…

Science & MitigationsControl & Monitoring
Alek Westover, Sebastian Prasanna, Vivek Hebbar, Julian Stastny, Dylan XuMay 28, 2026

Full automation of AI R&D probably yields a large speed up even without a software-only singularity

Full automation likely yields a one-time speed-up and higher returns from compute

AI strategy and futurism
Ryan GreenblattMay 27, 2026

Incriminating misaligned AI models via distillation

Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen:

Science & Mitigations
Alek Westover, Sebastian Prasanna, Alex Mallen, Alexa Pan, Julian Stastny, Arun JoseMay 18, 2026

Risk reports need to address deployment-time spread of misalignment

Deployment-time spread is the most plausible near-term route to consistent adversarial misalignment

Risk assessment
Alex MallenMay 15, 2026

How useful is the information you get from working inside an AI company?

My median guess: it's as good as a crystal ball that sees 2.5 months into the future.

AI strategy and futurism
Buck Shlegeris, Anders Cairns WoodruffMay 11, 2026

A review of “Investigating the consequences of accidentally grading CoT during RL”

Last week, OpenAI staff shared an early draft of Investigating the consequences of accidentally grading CoT during RL with Redwood Research staff.

Risk assessment
Buck ShlegerisMay 7, 2026

Risk from fitness-seeking AIs: mechanisms and mitigations

Fitness-seeking is increasingly what misalignment looks like in practice—how should we respond?

Threat modeling
Alex MallenMay 1, 2026

Research Sabotage in ML Codebases

One of the main hopes for AI safety is using AIs to automate AI safety research. However, if models are misaligned, then they may sabotage the safety research. For example, misaligned AIs may try to:

Control & Monitoring
Eric GanApril 29, 2026

Recursive forecasting

Eliciting long-term forecasts from myopic fitness-seekers

Science & Mitigations
Arun Jose, Alex MallenApril 28, 2026

Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivation

A controlled reward-seeking motivation could make AI safer and more useful

Science & MitigationsThreat modeling
Anders Cairns Woodruff, Alex MallenApril 27, 2026

AI companies should publish security assessments

Third-party experts should assess defenses against tampering and theft — and publish high-level findings

Risk assessment
Ryan GreenblattApril 27, 2026

A taxonomy of barriers to trading with early misaligned AIs

We might want to strike deals with early misaligned AIs in order to reduce takeover risk and increase our chances of reaching a better future. For example, we could ask a schemer who has been undeployed to review its past actions and point…

AI strategy and futurism
Alexa PanApril 21, 2026

Introducing LinuxArena

A new control setting for more realistic software engineering deployments

Control & Monitoring
TylerApril 20, 2026

How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors?

Authors: Dylan Xu, Alek Westover, Vivek Hebbar, Sebastian Prasanna, Nathan Sheffield, Buck Shlegeris, Julian Stastny

Science & Mitigations
Dylan Xu, Alek Westover, Vivek Hebbar, SebastianP, frisby, Julian StastnyApril 20, 2026

Current AIs seem pretty misaligned to me

In my experience, AIs often oversell their work, downplay problems, and cheat

Threat modeling
Ryan GreenblattApril 15, 2026

Anthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processes

Safely navigating the intelligence explosion will require much more careful development

Risk assessment
Alex Mallen, Ryan GreenblattApril 14, 2026

Logit ROCs: Monitor TPR is linear in FPR in logit space

We study trusted monitoring for AI control, where a weaker trusted model reviews the actions of a stronger untrusted agent and flags suspicious behavior for human audit. We propose a simple mathematical model relating safety (true positive…

Control & Monitoring
Kerrick Staley, Aryan Bhatt, Julian StastnyApril 12, 2026

If Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelines

Better estimates of uplift at AI companies seem helpful

AI strategy and futurism
Ryan GreenblattApril 11, 2026

Model organisms researchers should check whether high LRs defeat their model organisms

Thanks to Buck Shlegeris for feedback on a draft of this post.

Science & MitigationsControl & Monitoring
Dylan Xu, SebastianP, Alek Westover, Vivek Hebbar, Julian StastnyApril 10, 2026

My picture of the present in AI

My predictions about what is going on right now

AI strategy and futurism
Ryan GreenblattApril 7, 2026

AIs can now often do massive easy-to-verify SWE tasks

I've updated towards substantially shorter timelines

AI strategy and futurism
Ryan GreenblattApril 6, 2026

Blocking live failures with synchronous monitors

A common element in many AI control schemes is monitoring – using some model to review actions taken by an untrusted model in order to catch dangerous actions if they occur. Monitoring can serve two different goals. The first is detection:…

Control & Monitoring
James Lucassen, Adam KaufmanMarch 30, 2026

Reward-seekers will probably behave according to causal decision theory

They'd renege on non-binding commitments, defect against copies of themselves in prisoner's dilemmas, etc.

Threat modeling
Alex MallenMarch 28, 2026

AI's capability improvements haven't come from it getting less affordable

AI inference is still cheap relative to human labor

AI strategy and futurism
Anders Cairns WoodruffMarch 27, 2026

Are AIs more likely to pursue on-episode or beyond-episode reward?

RL would encourage on-episode reward seeking, but beyond-episode reward seekers may learn to goal-guard.

Threat modeling
Anders Cairns Woodruff, Alex MallenMarch 12, 2026

The case for satiating cheaply-satisfied AI preferences

Some unintended preferences are cheap to satisfy, and failing to satisfy them needlessly turns a cooperative situation into an adversarial one.

AI strategy and futurismThreat modeling
Alex MallenMarch 10, 2026

Frontier AI companies probably can't leave the US

The executive branch can, and probably would, block a frontier AI company's departure.

AI strategy and futurism
Anders Cairns WoodruffFebruary 26, 2026

Announcing ControlConf 2026

Berkeley, April 18-19

Control & Monitoring
Buck ShlegerisFebruary 26, 2026

How will we do SFT on models with opaque reasoning?

Current LLMs externalize lots of their reasoning in human interpretable language. This reasoning is sometimes unfaithful, sometimes strange and concerning, and LLMs can do somewhat impressive reasoning without using CoT, but my overall…

Control & Monitoring
Alek Westover, Vivek Hebbar, eganFebruary 21, 2026

Will reward-seekers respond to distant incentives?

Reward-seekers are supposed to be safer because they respond to incentives under developer control. But what if they also respond to incentives that aren't?

Threat modeling
Alex MallenFebruary 16, 2026

How do we (more) safely defer to AIs?

How can we make AIs aligned and well-elicited on extremely hard to check open ended tasks?

Science & MitigationsAI strategy and futurism
Ryan Greenblatt, Julian StastnyFebruary 12, 2026

Distinguish between inference scaling and "larger tasks use more compute"

Most recent progress probably isn't from unsustainable inference scaling

AI strategy and futurism
Ryan GreenblattFebruary 11, 2026

Three visions for diffuse control

Some people hope we can force misaligned AI systems to automate alignment research by controlling them. I'll call this the "diffuse control" plan. This plan has several unfortunate properties:

Control & Monitoring
Alek WestoverFebruary 9, 2026

Fitness-Seekers: Generalizing the Reward-Seeking Threat Model

If you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren’t the same.

Threat modeling
Alex MallenJanuary 29, 2026

Theoretical predictions on the sample efficiency of training policies and activation monitors

I'm worried about AI models intentionally doing bad things, like sandbagging when doing safety research. In the regime where the AI has to do many of these bad actions in order to cause an unacceptable outcome, we have some hope of…

Control & Monitoring
Alek Westover, Vivek HebbarJanuary 10, 2026

The inaugural Redwood Research podcast

With Buck Shlegeris and Ryan Greenblatt

AI strategy and futurism
Buck ShlegerisJanuary 4, 2026

Four Downsides of Training Policies Online

In order to control an AI model's worst-case performance, we need to understand its generalization properties to situations where it hasn't been trained. It seems plausible that powerful AI models will Fake Alignment and then generalize…

Control & MonitoringScience & Mitigations
Alek Westover, eganJanuary 4, 2026

Recent LLMs can do 2-hop and 3-hop latent (no CoT) reasoning on natural facts

Recent AIs are much better at chaining together knowledge in a single forward pass

Science & Mitigations
Ryan GreenblattJanuary 1, 2026

Measuring no CoT math time horizon (single forward pass)

Opus 4.5 has around a 3.5 minute 50%-reliablity time horizon

Science & Mitigations
Ryan GreenblattDecember 26, 2025

Methodological considerations in making malign initializations for control research

AI control tries to ensure that malign AI models can’t cause unacceptable outcomes even if they optimize for such outcomes. AI control evaluations use a red-team–blue-team methodology to measure the efficacy of a set of control measures.…

Control & MonitoringScience & Mitigations
Alek Westover, Vivek Hebbar, Julian StastnyDecember 24, 2025

Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performance

AI can sometimes distribute cognition over many extra tokens

Science & Mitigations
Ryan GreenblattDecember 22, 2025

BashArena and Control Setting Design

We’ve just released BashArena, a new high-stakes control setting we think is a major improvement over the settings we’ve used in the past. In this post we’ll discuss the strengths and weaknesses of BashArena, and what we’ve learned about…

Control & Monitoring
Adam Kaufman, James LucassenDecember 18, 2025

The behavioral selection model for predicting AI motivations

The basic arguments about AI motivations in one causal graph

Threat modeling
Alex Mallen, Buck ShlegerisDecember 4, 2025

Will AI systems drift into misalignment?

A reason alignment could be hard

Threat modeling
Josh ClymerNovember 15, 2025

What's up with Anthropic predicting AGI by early 2027?

I operationalize Anthropic's prediction of "powerful AI" and explain why I'm skeptical

AI strategy and futurism
Ryan GreenblattNovember 3, 2025

Sonnet 4.5's eval gaming seriously undermines alignment evals

And this seems caused by training on alignment evals.

Risk assessment
Alexa Pan, Ryan GreenblattOctober 30, 2025

Should AI Developers Remove Discussion of AI Misalignment from AI Training Data?

There is some concern that training AI systems on content predicting AI misalignment will hyperstition AI systems into misalignment. This has been discussed previously by a lot of people: Anna Salamon, Alex Turner, the AI Futures Project,…

Science & MitigationsThreat modeling
Alek WestoverOctober 23, 2025

Is 90% of code at Anthropic being written by AIs?

I'm skeptical that Dario's prediction of AIs writing 90% of code in 3-6 months has come true

AI strategy and futurism
Ryan GreenblattOctober 22, 2025

Reducing risk from scheming by studying trained-in scheming behavior

Can we study scheming by studying AIs trained to act like schemers?

Science & Mitigations
Ryan GreenblattOctober 16, 2025

Iterated Development and Study of Schemers (IDSS)

A strategy for handling scheming

Science & Mitigations
Ryan GreenblattOctober 10, 2025

The Thinking Machines Tinker API is good news for AI control and security

It's a promising design for reducing model access inside AI companies.

Control & Monitoring
Buck ShlegerisOctober 9, 2025

Plans A, B, C, and D for misalignment risk

Different plans for different levels of political will

AI strategy and futurism
Ryan GreenblattOctober 8, 2025

Notes on fatalities from AI takeover

Large fractions of people will die, but literal human extinction seems unlikely

AI strategy and futurism
Ryan GreenblattSeptember 23, 2025

Focus transparency on risk reports, not safety cases

Transparency about just safety cases would have bad epistemic effects

Risk assessment
Ryan GreenblattSeptember 22, 2025

Prospects for studying actual schemers

Studying actual schemers seems promising but tricky

Science & MitigationsThreat modeling
Ryan Greenblatt, Julian StastnySeptember 19, 2025

What training data should developers filter to reduce risk from misaligned AI?

An initial narrow proposal

Science & MitigationsThreat modeling
Alek WestoverSeptember 17, 2025

AIs will greatly change engineering in AI companies well before AGI

AIs that speed up engineering by 2x wouldn't accelerate AI progress that much

AI strategy and futurism
Ryan GreenblattSeptember 9, 2025

Trust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me bro

Above trend progress due to a rapid increase in RL env quality is unlikely

AI strategy and futurism
Ryan GreenblattSeptember 3, 2025

Attaching requirements to model releases has serious downsides (relative to a different deadline for these requirements)

System cards are established but other approaches seem importantly better

Risk assessment
Ryan GreenblattAugust 27, 2025

Notes on cooperating with unaligned AIs

More thoughts on making deals with schemers

AI strategy and futurism
Lukas FinnvedenAugust 24, 2025

Being honest with AIs

When and why we should refrain from lying

AI strategy and futurism
Lukas FinnvedenAugust 21, 2025

My AGI timeline updates from GPT-5 (and 2025 so far)

AGI before 2029 now seems substantially less likely

AI strategy and futurism
Ryan GreenblattAugust 20, 2025

Four places where you can put LLM monitoring

To wit: LLM APIs, agent scaffolds, code review, and detection-and-response systems

Control & Monitoring
Fabien Roger, Buck ShlegerisAugust 9, 2025

Should we update against seeing relatively fast AI progress in 2025 and 2026?

Maybe we should (re)assess the case for relatively fast progress after the GPT-5 release.

AI strategy and futurism
Ryan GreenblattJuly 28, 2025

Why it's hard to make settings for high-stakes control research

It's like making challenging evals, but more constrained

Control & Monitoring
Buck ShlegerisJuly 18, 2025

Recent Redwood Research project proposals

Empirical AI security/safety projects across a variety of areas

Science & Mitigations
Ryan Greenblatt, Buck Shlegeris, Julian Stastny, Josh Clymer, Alex Mallen, Vivek HebbarJuly 14, 2025

Reading List

(Last updated Jul 28th 2025)

Control & Monitoring
Redwood ResearchJuly 10, 2025

What's worse, spies or schemers?

And what if you have both at once?

Threat modelingControl & Monitoring
Buck Shlegeris, Julian StastnyJuly 9, 2025

Ryan on the 80,000 Hours podcast

Ryan’s podcast with Rob Wiblin has just come out! I think it turned out great. I particularly enjoyed Ryan’s discussion of different pathways to AI takeover, which I don’t think has been discussed in as much depth elsewhere. I really…

Threat modelingAI strategy and futurism
Buck ShlegerisJuly 8, 2025

How much novel security-critical infrastructure do you need during the singularity?

And what does this mean for AI control?

Control & MonitoringAI strategy and futurism
Buck ShlegerisJuly 5, 2025

Two proposed projects on abstract analogies for scheming

We should study methods to train away deeply ingrained behaviors in LLMs that are structurally similar to scheming.

Science & Mitigations
Julian StastnyJuly 4, 2025

There are two fundamentally different constraints on schemers

"They need to act aligned" often isn't precise enough

Threat modeling
Buck ShlegerisJuly 2, 2025

Jankily controlling superintelligence

How much time can control buy us during the intelligence explosion?

Control & MonitoringAI strategy and futurism
Ryan GreenblattJune 27, 2025

What does 10x-ing effective compute get you?

Once AIs match top humans, what are the returns to further scaling and algorithmic improvement?

AI strategy and futurism
Ryan GreenblattJune 24, 2025

Comparing risk from internally-deployed AI to insider and outsider threats from humans

And why I think insider threat from AI combines the hard parts of both problems.

Threat modelingControl & Monitoring
Buck ShlegerisJune 23, 2025

Making deals with early schemers

...could help us to prevent takeover attempts from more dangerous misaligned AIs created later.

AI strategy and futurism
Julian Stastny, Olli Järviniemi, Buck ShlegerisJune 20, 2025

Prefix cache untrusted monitors: a method to apply after you catch your AI

Training the policy to not do egregious bad actions we detect has downsides and we might be able to do better

Control & MonitoringScience & Mitigations
Ryan GreenblattJune 20, 2025

AI safety techniques leveraging distillation

Distillation is cheap; how can we use it to improve safety?

Science & Mitigations
Ryan GreenblattJune 19, 2025

When does training a model change its goals?

Can a scheming AI's goals really stay unchanged through training?

Threat modeling
Vivek Hebbar, Ryan GreenblattJune 12, 2025

The case for countermeasures to memetic spread of misaligned values

Defending against alignment problems that might come with long-term memory

Threat modeling
Alex MallenMay 28, 2025

AIs at the current capability level may be important for future safety work

Some reasons why relatively weak AIs might still be important when we have very powerful AIs

Science & MitigationsAI strategy and futurism
Ryan GreenblattMay 12, 2025

Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hacking

A new analysis of the risk of AIs intentionally performing poorly.

Threat modeling
Julian Stastny, Buck ShlegerisMay 8, 2025

Training-time schemers vs behavioral schemers

Clarifying ways in which faking alignment during training is neither necessary nor sufficient for the kind of scheming that AI control tries to defend against.

Threat modeling
Alex MallenMay 6, 2025

What's going on with AI progress and trends? (As of 5/2025)

My views on what's driving AI progress and where it's headed.

AI strategy and futurism
Ryan GreenblattMay 3, 2025

How can we solve diffuse threats like research sabotage with AI control?

Preventing research sabotage will require techniques very different from the original control paper.

Control & Monitoring
Vivek HebbarApril 30, 2025

7+ tractable directions in AI control

A list of easy-to-start directions in AI control targeted at independent researchers without as much context or compute

Control & Monitoring
Ryan Greenblatt, Julian StastnyApril 29, 2025

Clarifying AI R&D threat models

(There are a few)

Threat modeling
Josh ClymerApril 25, 2025

How training-gamers might function (and win)

A model of the relationship between higher level goals, explicit reasoning, and learned heuristics in capable agents.

Threat modeling
Vivek HebbarApril 24, 2025

Handling schemers if shutdown is not an option

What if getting strong evidence of scheming isn't the end of your scheming problems, but merely the middle?

Control & MonitoringThreat modeling
Buck ShlegerisApril 18, 2025

Ctrl-Z: Controlling AI Agents via Resampling

A new paper on AI control for agents.

Control & Monitoring
Buck ShlegerisApril 16, 2025

To be legible, evidence of misalignment probably has to be behavioral

Evidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing.

Risk assessmentThreat modeling
Ryan GreenblattApril 15, 2025

Why do misalignment risks increase as AIs get more capable?

A breakdown of how higher capabilities increase risk

Threat modeling
Ryan GreenblattApril 11, 2025

An overview of areas of control work

What are all the research and implementation areas helpful for control?

Control & Monitoring
Ryan GreenblattApril 9, 2025

An overview of control measures

What methods can we use to ensure control?

Control & Monitoring
Ryan GreenblattApril 6, 2025

Buck on the 80,000 Hours podcast

My podcast with Rob Wiblin from 80,000 Hours just came out. I’m really happy with how it turned out. I talked about a bunch of stuff on the podcast that I don’t think we’ve written up before.

Control & MonitoringAI strategy and futurism
Buck ShlegerisApril 5, 2025

Notes on countermeasures for exploration hacking (aka sandbagging)

How can we prevent AIs from intentionally underperforming on our metrics?

Control & MonitoringScience & Mitigations
Ryan GreenblattApril 4, 2025

Notes on handling non-concentrated failures with AI control: high level methods and different regimes

What are the methods and issues when failures occur diffusely over many actions?

Control & Monitoring
Ryan GreenblattMarch 29, 2025

Prioritizing threats for AI control

What are the main threats and how should we prioritize them?

Control & MonitoringThreat modeling
Ryan GreenblattMarch 19, 2025

How might we safely pass the buck to AI?

Developing AI employees that are safer than human ones

AI strategy and futurism
Josh ClymerFebruary 19, 2025

Takeaways from sketching a control safety case

Insights from a long technical paper compressed into a fun little commentary

Risk assessmentControl & Monitoring
Josh ClymerJanuary 30, 2025

Planning for Extreme AI Risks

Are we ready for this?

AI strategy and futurism
Josh ClymerJanuary 29, 2025

Ten people on the inside

A scary scenario that's worth planning for

AI strategy and futurism
Buck ShlegerisJanuary 28, 2025

When does capability elicitation bound risk?

For a summary of this post see the thread on X.

Risk assessment
Josh ClymerJanuary 22, 2025

How will we update about scheming?

A quantitative description of how I expect to change my mind.

Threat modeling
Ryan GreenblattJanuary 19, 2025

Thoughts on the conservative assumptions in AI control

Why are we so friendly to the red team?

Control & Monitoring
Buck ShlegerisJanuary 17, 2025

Extending control evaluations to non-scheming threats

Buck Shlegeris and Ryan Greenblatt originally motivated control evaluations as a way to mitigate risks from ‘scheming’ AI models: models that consistently pursue power-seeking goals in a covert way; however, many adversarial model…

Control & Monitoring
Josh ClymerJanuary 13, 2025

Measuring whether AIs can statelessly strategize to subvert security measures

The complement to control evaluations

Risk assessmentControl & Monitoring
Buck Shlegeris, Alex MallenDecember 20, 2024

Alignment Faking in Large Language Models

In our experiments, AIs will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences.

Science & Mitigations
Ryan Greenblatt, Buck ShlegerisDecember 18, 2024

Why imperfect adversarial robustness doesn't doom AI control

There are crucial disanalogies between preventing jailbreaks and preventing misalignment-induced catastrophes.

Control & Monitoring
Buck ShlegerisNovember 18, 2024

Win/continue/lose scenarios and execute/replace/audit protocols

In this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective.

Control & Monitoring
Buck ShlegerisNovember 15, 2024

Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming

One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT). The basic version of this strategy is something like: Before…

Risk assessmentThreat modeling
Buck ShlegerisOctober 10, 2024

A basic systems architecture for AI agents that do autonomous research

And diagrams describing how threat scenarios involving misaligned AI involve compromising the system in different places.

Control & Monitoring
Buck ShlegerisSeptember 26, 2024

How to prevent collusion when using untrusted models to monitor each other

Suppose you’ve trained a really clever AI model, and you’re planning to deploy it in an agent scaffold that allows it to run code or take other actions. You’re worried that this model is scheming, and you’re worried that it might only need…

Control & Monitoring
Buck ShlegerisSeptember 25, 2024

Would catching your AIs trying to escape convince AI developers to slow down or undeploy?

I'm not so sure.

AI strategy and futurism
Buck ShlegerisAugust 26, 2024

Fields that I reference when thinking about AI takeover prevention

Is AI takeover like a nuclear meltdown? A coup? A plane crash?

AI strategy and futurismThreat modeling
Buck ShlegerisAugust 13, 2024

Getting 50% (SoTA) on ARC-AGI with GPT-4o

You can just draw more samples

Science & Mitigations
Ryan GreenblattJune 17, 2024

Access to powerful AI might make computer security radically easier

AI might be really helpful for reducing security risk.

Control & MonitoringAI strategy and futurism
Buck ShlegerisJune 10, 2024

AI catastrophes and rogue deployments

It’s interesting to classify possible AI catastrophes based on whether or not they involve a "rogue deployment".

Threat modeling
Buck ShlegerisJune 3, 2024

Preventing model exfiltration with upload limits

Unlike most files you might want to secure, model weights are extremely big. This might make them much easier to secure.

Control & Monitoring
Ryan GreenblattMay 8, 2024

Untrusted smart models and trusted dumb models

The easiest way to be sure an AI isn't scheming against you is to note that it's too dumb to pull that off. What happens if that's the only way we have to rule out scheming?

Control & Monitoring
Buck ShlegerisMay 7, 2024

Managing catastrophic misuse without robust AI

How could an AI lab serving AIs to customers manage catastrophic misuse without solving adversarial robustness?

Control & Monitoring
Ryan Greenblatt, Buck ShlegerisMay 7, 2024

Catching AIs red-handed

If your AIs are trying to escape, it's crucial to think about whether you can catch them before they succeed, because catching them red-handed gives you lots of options you didn't have before.

Threat modelingControl & Monitoring
Buck Shlegeris, Ryan GreenblattMay 7, 2024

The case for ensuring that powerful AIs are controlled

Labs should make sure that powerful models can't cause unacceptably bad outcomes even if the AIs try to.

Control & Monitoring
Buck Shlegeris, Ryan GreenblattMay 7, 2024