“Iterated Development and Study of Schemers (IDSS)” by ryan_greenblatt

Update: 2025-10-10

Description

In a previous post, we discussed prospects for studying scheming using natural examples. In this post, we'll describe a more detailed proposal for iteratively constructing scheming models, techniques for detecting scheming, and techniques for preventing scheming. We'll call this strategy Iterated Development and Study of Schemers (IDSS). We'll be using concepts from that prior post, like the idea of trying to make schemers which are easier to catch.

Two key difficulties with using natural examples of scheming are that it is hard to catch (and re-catch) schemers and that it's hard to (cheaply) get a large number of diverse examples of scheming to experiment on. One approach for partially resolving these issues is to experiment on weak schemers which are easier to catch and cheaper to experiment on. However, these weak schemers might not be analogous to the powerful schemers which are actually dangerous, and these weak AIs [...]

---

Outline:

(01:36 ) The IDSS strategy

(02:10 ) Iterated development of scheming testbeds

(04:35 ) Gradual addition of capabilities

(08:39 ) Using improved testbeds for improving scheming mitigation techniques

(09:09 ) What can go wrong with IDSS?

(13:20 ) Conclusion

The original text contained 2 footnotes which were omitted from this narration.

---

First published:

October 10th, 2025

Source:

https://www.lesswrong.com/posts/QpzTmFLXMJcdRkPLZ/iterated-development-and-study-of-schemers-idss

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Comments

In Channel

“Emil the Moose” by Martin Sustrik

2025-10-1103:05

“Experiments With Sonnet 4.5 Fiction” by Tomás B.

2025-10-1111:06

“The Most Common Bad Argument In These Parts” by J Bostock

2025-10-1108:12

“Iterated Development and Study of Schemers (IDSS)” by ryan_greenblatt

2025-10-1014:20

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

2025-10-1007:41

“‘Yes, and—’ Requires the Possibility of ‘No, Because—’” by Zack_M_Davis

2025-10-1006:42

“Stars are a rounding error” by Algon

2025-10-1005:48

“Training Qwen-1.5B with a CoT legibility penalty” by Fabien Roger

2025-10-1010:20

“At odds with the unavoidable meta-message” by Ruby

2025-10-1007:28

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

2025-10-0917:35

“I take antidepressants. You’re welcome” by Elizabeth

2025-10-0906:10

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

2025-10-0921:54

“The Thinking Machines Tinker API is good news for AI control and security” by Buck

2025-10-0911:54

“Hospitalization: A Review” by Logan Riggs

2025-10-0918:53

“The Relationship Between Social Punishment and Shared Maps” by Zack_M_Davis

2025-10-0908:18

“Spooky Collusion at a Distance with Superrational AI” by bira

2025-10-0913:14

“Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior” by Sam Marks

2025-10-0804:07

“Plans A, B, C, and D for misalignment risk” by ryan_greenblatt

2025-10-0812:02

“Irresponsible Companies Can Be Made of Responsible Employees” by VojtaKovarik

2025-10-0809:33

“Replacing RL w/ Parameter-based Evolutionary Strategies” by Logan Riggs

2025-10-0808:30

00:00

“Iterated Development and Study of Schemers (IDSS)” by ryan_greenblatt

#box-pro-ellipsis-176026416527174{-webkit-line-clamp:2;}“Iterated Development and Study of Schemers (IDSS)” by ryan_greenblatt

“Iterated Development and Study of Schemers (IDSS)” by ryan_greenblatt

“Iterated Development and Study of Schemers (IDSS)” by ryan_greenblatt