Listen Top Shows Blog

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

Update: 2025-10-09

Share

Description

TL;DR: I made a dataset of realistic harmless reward hacks and fine-tuned GPT-4.1 on it. The resulting models don't show emergent misalignment on the standard evals, but they do alignment fake (unlike models trained on toy reward hacks), seem more competently misaligned, are highly evaluation-aware, and the effects persist when mixing in normal data.

Thanks to Aidan Ewart, Jack Kaunismaa, Abhay Sheshadri, Maxime Riché, Axel Ahlqvist, Niels Warncke, Daniel Tan, Carolyn Qian, and Kei Nishimura-Gasparian for helpful conversations, comments and/or feedback. This post is best viewed as an informal report on preliminary results done over a couple days, rather than a very polished analysis.

Introduction

Taylor et al finds that fine-tuning LLMs on harmless reward hacks causes generalization to unrelated misaligned behavior on the emergent misalignment (EM) evals. They constructed a fine-tuning dataset (School of Reward Hacks) of samples like this:

There's a details box here with the title "Sample [...]

---

Outline:

(00:56 ) Introduction

(03:17 ) Dataset

(05:24 ) Emergent Misalignment Evals

(07:34 ) Alignment Faking

(16:29 ) Takeaways

(18:28 ) How robust is this effect?

The original text contained 11 footnotes which were omitted from this narration.

---

First published:

October 9th, 2025

Source:

https://www.lesswrong.com/posts/HLJoJYi52mxgomujc/realistic-reward-hacking-induces-different-and-deeper-1

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Comments

In Channel

“Experiments With Sonnet 4.5 Fiction” by Tomás B.

“Experiments With Sonnet 4.5 Fiction” by Tomás B.

2025-10-1111:06

“The Most Common Bad Argument In These Parts” by J Bostock

“The Most Common Bad Argument In These Parts” by J Bostock

2025-10-1108:12

“Iterated Development and Study of Schemers (IDSS)” by ryan_greenblatt

“Iterated Development and Study of Schemers (IDSS)” by ryan_greenblatt

2025-10-1014:20

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

“Training fails to elicit subtle reasoning in current language models” by mishajw, Fabien Roger, Hoagy, gasteigerjo, Joe Benton, Vlad Mikulik

2025-10-1007:41

“‘Yes, and—’ Requires the Possibility of ‘No, Because—’” by Zack_M_Davis

“‘Yes, and—’ Requires the Possibility of ‘No, Because—’” by Zack_M_Davis

2025-10-1006:42

“Stars are a rounding error” by Algon

“Stars are a rounding error” by Algon

2025-10-1005:48

“Training Qwen-1.5B with a CoT legibility penalty” by Fabien Roger

“Training Qwen-1.5B with a CoT legibility penalty” by Fabien Roger

2025-10-1010:20

“At odds with the unavoidable meta-message” by Ruby

“At odds with the unavoidable meta-message” by Ruby

2025-10-1007:28

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

“Towards a Typology of Strange LLM Chains-of-Thought” by 1a3orn

2025-10-0917:35

“I take antidepressants. You’re welcome” by Elizabeth

“I take antidepressants. You’re welcome” by Elizabeth

2025-10-0906:10

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

2025-10-0921:54

“The Thinking Machines Tinker API is good news for AI control and security” by Buck

“The Thinking Machines Tinker API is good news for AI control and security” by Buck

2025-10-0911:54

“Hospitalization: A Review” by Logan Riggs

“Hospitalization: A Review” by Logan Riggs

2025-10-0918:53

“The Relationship Between Social Punishment and Shared Maps” by Zack_M_Davis

“The Relationship Between Social Punishment and Shared Maps” by Zack_M_Davis

2025-10-0908:18

“Spooky Collusion at a Distance with Superrational AI” by bira

“Spooky Collusion at a Distance with Superrational AI” by bira

2025-10-0913:14

“Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior” by Sam Marks

“Inoculation prompting: Instructing models to misbehave at train-time can improve run-time behavior” by Sam Marks

2025-10-0804:07

“Plans A, B, C, and D for misalignment risk” by ryan_greenblatt

“Plans A, B, C, and D for misalignment risk” by ryan_greenblatt

2025-10-0812:02

“Irresponsible Companies Can Be Made of Responsible Employees” by VojtaKovarik

“Irresponsible Companies Can Be Made of Responsible Employees” by VojtaKovarik

2025-10-0809:33

“Replacing RL w/ Parameter-based Evolutionary Strategies” by Logan Riggs

“Replacing RL w/ Parameter-based Evolutionary Strategies” by Logan Riggs

2025-10-0808:30

“You Should Get a Reusable Mask” by jefftk

“You Should Get a Reusable Mask” by jefftk

2025-10-0803:10

00:00

00:00

1.0x

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien

“Realistic Reward Hacking Induces Different and Deeper Misalignment” by Jozdien