“Serious Flaws in CAST” by Max Harms

Update: 2025-11-19

Description

Last year I wrote the CAST agenda, arguing that aiming for Corrigibility As Singular Target was the least-doomed way to make an AGI. (Though it is almost certainly wiser to hold off on building it until we have more skill at alignment, as a species.)

I still basically believe that CAST is right. Corrigibility still seems like a promising target compared to full alignment with human values, since there's a better story for how a near-miss when aiming towards corrigibility might be recoverable, but a near-miss when aiming for goodness could result is a catastrophe, due to the fragility of value. On top of this, corrigibility is significantly simpler and less philosophically fraught than human values, decreasing the amount of information that needs to be perfectly transmitted to the machine. Any spec, constitution, or whatever that attempts to balance corrigibility with other goals runs the risk of the convergent instrumental drives towards those other goals washing out the corrigibility. My most recent novel is intended to be an introduction to corrigibility that's accessible to laypeople, featuring a CAST AGI as a main character, and I feel good about what I wrote there.

But I'm starting to feel like certain [...]

---

Outline:

(02:42 ) 1. Oops I Ruined the Universe

(06:12 ) Is There an Obvious Fix?

(08:20 ) 2. Attractor Basin Masks Brittleness

(11:32 ) 3. Feedback Loops Quickly Disappear by Default

The original text contained 2 footnotes which were omitted from this narration.

---

First published:

November 19th, 2025

Source:

https://www.lesswrong.com/posts/qgBFJ72tahLo5hzqy/serious-flaws-in-cast

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Comments

In Channel

“Preventing covert ASI development in countries within our agreement” by Aaron_Scher

2025-11-2022:47

“Current LLMs seem to rarely detect CoT tampering” by Bart Bussmann, Arthur Conmy, Neel Nanda, Senthooran Rajamanoharan, Josh Engels, Bartosz Cywiński

2025-11-1916:18

“The Bughouse Effect” by TsviBT

2025-11-1927:18

“Serious Flaws in CAST” by Max Harms

2025-11-1914:41

“Memories of a British Boarding School #2” by Ben Pace

2025-11-1912:41

“Automate, automate it all” by habryka

2025-11-1909:25

“How the aliens next door shower” by Ruby

2025-11-1906:02

“Victor Taelin’s notes on Gemini 3” by Gunnar_Zarncke

2025-11-1906:45

“Anthropic is (probably) not meeting its RSP security commitments” by habryka

2025-11-1908:58

“Considerations for setting the FLOP thresholds in our example international AI agreement” by peterbarnett, Aaron_Scher

2025-11-1914:28

“On Writing #2” by Zvi

2025-11-1824:04

“New Report: An International Agreement to Prevent the Premature Creation of Artificial Superintelligence” by Aaron_Scher, David Abecassis, Brian Abeyta, peterbarnett

2025-11-1806:53

“Status Is The Game Of The Losers’ Bracket” by johnswentworth

2025-11-1808:01

“Eat The Richtext” by dreeves

2025-11-1804:51

“Small batches and the mythical single piece flow” by habryka

2025-11-1809:47

“How Colds Spread” by RobertM

2025-11-1820:32

“Middlemen Are Eating the World (And That’s Good, Actually)” by Linch

2025-11-1808:52

“Why is American mass-market tea so terrible?” by RobertM

2025-11-1805:05

“An Analogue Of Set Relationships For Distribution” by johnswentworth, David Lorell

2025-11-1808:37

“AI 2025 - Last Shipmas” by Simon Lermen

2025-11-1819:07

00:00

“Serious Flaws in CAST” by Max Harms

#box-pro-ellipsis-176361144453345{-webkit-line-clamp:2;}“Serious Flaws in CAST” by Max Harms

“Serious Flaws in CAST” by Max Harms

“Serious Flaws in CAST” by Max Harms