“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

Update: 2025-10-10

Description

While recent AI systems achieve strong performance through human-readable reasoning that should be simple to monitor (OpenAI, 2024, Anthropic, 2025), we investigate whether models can learn to reason about malicious side tasks while making that reasoning appear benign. We find that Sonnet 3.7 can learn to evade either a reasoning monitor, by persuading the monitor that a blatant backdoor is benign, or an output-only monitor, by devising sophisticated backdoors that the output-only monitor doesn’t detect. But when trained to evade both reasoning and output-only monitors, Sonnet 3.7 is unable to use reasoning to improve its backdoor success rate without triggering a reasoning monitor. Like previous work (Baker et al., 2025, Emmons et al., 2025), our results suggest that reasoning monitors can provide strong assurance that language models are not pursuing reasoning-heavy malign side tasks, but that additional mitigations may be required for robustness to monitor persuasion.

Figure 1: We trained [...]

---

First published:

October 9th, 2025

Source:

https://www.lesswrong.com/posts/MmuyzfsaNrSvRCsFk/training-fails-to-elicit-subtle-reasoning-in-current

---

Narrated by TYPE III AUDIO.

---

Images from the article:

$Figure 1: We trained three versions of Sonnet 3.7 to create subtle backdoors that circumvent different types of monitors: an output monitor (sees final code, comments stripped), a reasoning monitor (sees all reasoning including final output), and an output + reasoning monitor (logical OR of the previous two). For each variant, we measure backdoor success – the fraction of backdoors that are both functional and undetected by the monitor – conditioned on success in the main task. We found that: (a) training against the output monitor allowed the model to use reasoning to improve backdoor performance; (b) training against the reasoning monitor enabled the model to convince the monitor that malign backdoors were benign; and (c) when trained against both, the model failed to learn to use reasoning to improve performance.$

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Comments

In Channel

“Emil the Moose” by Martin Sustrik

2025-10-1103:05

“Experiments With Sonnet 4.5 Fiction” by Tomás B.

2025-10-1111:06

“The Most Common Bad Argument In These Parts” by J Bostock

2025-10-1108:12

“Iterated Development and Study of Schemers (IDSS)” by ryan_greenblatt

2025-10-1014:20

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

2025-10-1007:41

“‘Yes, and—’ Requires the Possibility of ‘No, Because—’” by Zack_M_Davis

2025-10-1006:42

“Stars are a rounding error” by Algon

2025-10-1005:48

“Training Qwen-1.5B with a CoT legibility penalty” by Fabien Roger

2025-10-1010:20

“At odds with the unavoidable meta-message” by Ruby

2025-10-1007:28

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

2025-10-0917:35

“I take antidepressants. You’re welcome” by Elizabeth

2025-10-0906:10

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

2025-10-0921:54

“The Thinking Machines Tinker API is good news for AI control and security” by Buck

2025-10-0911:54

“Hospitalization: A Review” by Logan Riggs

2025-10-0918:53

“The Relationship Between Social Punishment and Shared Maps” by Zack_M_Davis

2025-10-0908:18

“Spooky Collusion at a Distance with Superrational AI” by bira

2025-10-0913:14

“Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior” by Sam Marks

2025-10-0804:07

“Plans A, B, C, and D for misalignment risk” by ryan_greenblatt

2025-10-0812:02

“Irresponsible Companies Can Be Made of Responsible Employees” by VojtaKovarik

2025-10-0809:33

“Replacing RL w/ Parameter-based Evolutionary Strategies” by Logan Riggs

2025-10-0808:30

00:00

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

#box-pro-ellipsis-176028129317235{-webkit-line-clamp:2;}“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik