“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

Update: 2025-10-09

Description

Intro

LLMs being trained with RLVR (Reinforcement Learning from Verifiable Rewards) start off with a 'chain-of-thought' (CoT) in whatever language the LLM was originally trained on. But after a long period of training, the CoT sometimes starts to look very weird; to resemble no human language; or even to grow completely unintelligible.

Why might this happen?

I've seen a lot of speculation about why. But a lot of this speculation narrows too quickly, to just one or two hypotheses. My intent is also to speculate, but more broadly.

Specifically, I want to outline six nonexclusive possible causes for the weird tokens: new better language, spandrels, context refresh, deliberate obfuscation, natural drift, and conflicting shards.

And I also wish to extremely roughly outline ideas for experiments and evidence that could help us distinguish these causes.

I'm sure I'm not enumerating the full space of [...]

---

Outline:

(00:11 ) Intro

(01:34 ) 1. New Better Language

(04:06 ) 2. Spandrels

(06:42 ) 3. Context Refresh

(10:48 ) 4. Deliberate Obfuscation

(12:36 ) 5. Natural Drift

(13:42 ) 6. Conflicting Shards

(15:24 ) Conclusion

---

First published:

October 9th, 2025

Source:

https://www.lesswrong.com/posts/qgvSMwRrdqoDMJJnD/towards-a-typology-of-strange-llm-chains-of-thought

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Comments

In Channel

“Iterated Development and Study of Schemers (IDSS)” by ryan_greenblatt

2025-10-1014:20

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

2025-10-1007:41

“‘Yes, and—’ Requires the Possibility of ‘No, Because—’” by Zack_M_Davis

2025-10-1006:42

“Stars are a rounding error” by Algon

2025-10-1005:48

“Training Qwen-1.5B with a CoT legibility penalty” by Fabien Roger

2025-10-1010:20

“At odds with the unavoidable meta-message” by Ruby

2025-10-1007:28

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

2025-10-0917:35

“I take antidepressants. You’re welcome” by Elizabeth

2025-10-0906:10

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

2025-10-0921:54

“The Thinking Machines Tinker API is good news for AI control and security” by Buck

2025-10-0911:54

“Hospitalization: A Review” by Logan Riggs

2025-10-0918:53

“The Relationship Between Social Punishment and Shared Maps” by Zack_M_Davis

2025-10-0908:18

“Spooky Collusion at a Distance with Superrational AI” by bira

2025-10-0913:14

“Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior” by Sam Marks

2025-10-0804:07

“Plans A, B, C, and D for misalignment risk” by ryan_greenblatt

2025-10-0812:02

“Irresponsible Companies Can Be Made of Responsible Employees” by VojtaKovarik

2025-10-0809:33

“Replacing RL w/ Parameter-based Evolutionary Strategies” by Logan Riggs

2025-10-0808:30

“You Should Get a Reusable Mask” by jefftk

2025-10-0803:10

“Bending The Curve” by Zvi

2025-10-0740:12

[Linkpost] “Petri: An open-source auditing tool to accelerate AI safety research” by Sam Marks

2025-10-0703:33

00:00

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

#box-pro-ellipsis-176018441352055{-webkit-line-clamp:2;}“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn