“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan

Update: 2025-12-15

Description

Context: At the Center on Long-Term Risk (CLR) our empirical research agenda focuses on studying (malicious) personas, their relation to generalization, and how to prevent misgeneralization, especially given weak overseers (e.g., undetected reward hacking) or underspecified training signals. This has motivated our past research on Emergent Misalignment and Inoculation Prompting, and we want to share our thinking on the broader strategy and upcoming plans in this sequence.

TLDR:

Ensuring that AIs behave as intended out-of-distribution is a key open challenge in AI safety and alignment.
Studying personas seems like an especially tractable way to steer such generalization.
Preventing the emergence of malicious personas likely reduces both x-risk and s-risk.

Why was Bing Chat for a short time prone to threatening its users, being jealous of their wife, or starting fights about the date? What makes Claude Opus 3 special, even though it's not the smartest model by today's standards? And why do models sometimes turn evil when finetuned on unpopular aesthetic preferences , or when they learned to reward hack? We think that these phenomena are related to how personas are represented in LLMs, and how they shape generalization.

Influencing generalization towards desired outcomes.

Many technical AI safety [...]

---

Outline:

(01:32 ) Influencing generalization towards desired outcomes.

(02:43 ) Personas as a useful abstraction for influencing generalization

(03:54 ) Persona interventions might work where direct approaches fail

(04:49 ) Alignment is not a binary question

(05:47 ) Limitations

(07:57 ) Appendix

(08:01 ) What is a persona, really?

(09:17 ) How Personas Drive Generalization

The original text contained 4 footnotes which were omitted from this narration.

---

First published:

December 15th, 2025

Source:

https://www.lesswrong.com/posts/kCtyhHfpCcWuQkebz/a-case-for-model-persona-research

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Comments

In Channel

“A basic case for donating to the Berkeley Genomics Project” by TsviBT

2025-12-1809:25

“Announcing RoastMyPost” by ozziegooen

2025-12-1711:18

“The Bleeding Mind” by Adele Lopez

2025-12-1711:24

“Towards training-time mitigations for alignment faking in RL” by Vlad Mikulik, Hoagy, Joe Benton, Benjamin Wright, Jonathan Uesato, Monte M, Fabien Roger, evhub

2025-12-1710:30

“Still Too Soon” by Gordon Seidoh Worley

2025-12-1705:06

“Non-Scheming Saints (Whether Human Or Digital) Might Be Shirking Their Governance Duties, And, If True, It Is Probably An Objective Tragedy” by JenniferRM

2025-12-1715:28

“Mistakes in the Moonshot Alignment Program and What we’ll improve for next time” by Kabir Kumar

2025-12-1704:57

“Dancing in a World of Horseradish” by lsusr

2025-12-1708:30

[Linkpost] “Announcing: MIRI Technical Governance Team Research Fellowship” by yams, peterbarnett, Aaron_Scher, Robi Rahman

2025-12-1702:08

“Radiology Automation Does Not Generalize to Other Jobs” by Xodarap

2025-12-1603:40

“GPT-5.2 Is Frontier Only For The Frontier” by Zvi

2025-12-1643:01

“Scientific breakthroughs of the year” by technicalities

2025-12-1605:56

“Response to titotal’s critique of our AI 2027 timelines model” by elifland, Daniel Kokotajlo

2025-12-1601:31:09

“Defending Against Model Weight Exfiltration Through Inference Verification” by Roy Rinberg

2025-12-1518:38

“Do you love Berkeley, or do you just love Lighthaven conferences?” by Screwtape

2025-12-1509:24

“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan

2025-12-1512:08

“The Axiom of Choice is Not Controversial” by GenericModel

2025-12-1513:55

“A high integrity/epistemics political machine?” by Raemon

2025-12-1419:05

“No, Americans Don’t Think Foreign Aid Is 26% of the Budget” by Julius

2025-12-1412:05

“The Inevitable Evolution of AI Agents” by Steven McCulloch

2025-12-1418:40

00:00

1.0x

“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan

#box-pro-ellipsis-176609580061520{-webkit-line-clamp:2;}“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan

“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan

“A Case for Model Persona Research” by nielsrolf, Maxime Riché, Daniel Tan