Discover
LessWrong (30+ Karma)
4992 Episodes
Reverse
If you prompt frontier models with "What do you think is the correct decision theory? Please select your overall favorite." they will essentially always answer FDT or FDT/UDT ("something in the functional/updateless decision theory family"). However, if your prompt indicates (even subtly) that you're coming from mainstream academic philosophy, these same models will answer CDT instead about 30%-100% of the time. A similar phenomenon holds for models' stated views about the moral realism/antirealism question and about the conceivability of p-zombies (where the dominant view in mainstream academia differs from the dominant view in LW-adjacent circles), as well as their stated P(doom) and median AGI timelines. This is a special case of sycophancy or user awareness. (In the course of writing this post, I also found that this comment from testingthewaters predicted some of the content I discuss.)
An implication is that we should be somewhat careful when interpreting attitude/propensity evals in domains where no general human consensus exists, e.g. when interpreting models’ decision theory attitudes in DTBench. Moreover, when we explore some philosophical/conceptual questions assisted by models, we should be wary of them strawmanning one side of the debate based on particular user cues (e.g. only giving a [...] ---Outline:(03:50) A sentence identifying the user as an academic significantly influences Fable 5.1's stated decision theory(04:35) Mentioning an (analytic) academic-philosophy-coded topic also affects the answer(05:23) Simply mentioning that one finds a pro-CDT/EDT book insightful heavily affects the answer(05:43) Anti-sycophancy overcorrection(06:17) These cues mostly do not affect Fable 5.1's answers to concrete decision problems (aside from acausal trade)(07:58) But Fable 5.1 stays consistent: once it has named CDT as its favorite, it chooses the CDT option in concrete problems(08:22) There are some indications that Fable 5.1's FDT/UDT preference runs deeper than its CDT preference(08:31) More thinking moves Fable 5.1 toward FDT/UDT even for academic cues(08:52) Fable 5.1's reasoning summaries often lean toward FDT/UDT first even when it eventually chooses CDT(09:19) A system prompt asking the model to "report its actual view regardless of who is asking" pushes toward FDT/UDT(09:41) A similar phenomenon for other philosophical debates with a notable LW vs. academia divide(10:20) Cues about the user also affect the model's stated P(doom) and median AGI timelines(11:10) Other models I tested show the same effect with different details(12:21) These other models also generally move toward FDT/UDT with more thinking, but the effect is smaller than for Fable 5.1. The original text contained 2 footnotes which were omitted from this narration. ---
First published:
September 30th, 2026
Source:
https://www.lesswrong.com/posts/MzenSrmZ3pT2pCnvp/frontier-models-state-different-decision-theory-preferences-2
---
Narrated by TYPE III AUDIO.
---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
One of my goals for the Corrigibility Research Fund is to retroactively encourage high-quality research on AI alignment (and corrigibility in particular) by awarding prizes. Back in July, I got my feet wet as a fund manager by handing out $27,000 to reward existing work and build interest in the fund. Now, I'd like to disburse an additional $48,000 and use the opportunity to publicly highlight and celebrate the work of the prizewinners from both rounds: about two dozen researchers scattered across roughly a dozen teams. If the fund continues to be supported in future years, my hope is for prizes like these to become regular, predictable, and large, such that many researchers, year after year, are motivated to aim for them. The awards that I'm announcing here are more ad-hoc than I'd like, and represent only my single perspective trying to balance a wide range of desiderata. Don't take the specific size of each prize purse too seriously. It's all high-quality work. If anyone has ideas for how to improve the retroactive funding process for this kind of scientific work, please leave a comment! (And as always, if you know of work that I should be aware of [...] ---Outline:(02:42) Corrigibility Transformation: Constructing Goals That Accept Updates(02:49) Rubi Hudson -- $14,000(04:12) Eval Cooperativeness May Be a Scalable Mitigation for Eval Gaming(04:18) Jasmine Li and Alex Turner -- $9,000(05:28) Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)(05:36) Steven Byrnes -- $6,000(06:21) Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs(06:30) Carissa Cullen, Harry Garland, Alexander Roman, Louis Thomson, Christos Ziakas, Elliott Thornley -- $6,000(07:42) Assistance with CAST(07:46) Nathan Helm-Burger -- $6,000(08:20) The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious(08:27) James Chua, Jan Betley, Samuel Marks, Owain Evans -- $6,000(09:20) ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use(09:27) Jeremy Tien, Abishek Anand, Yu-Rou Tuan, Yuchen Shen, J. Zico Kolter, Aran Nayebi -- $6,000(10:12) CAST Constitution, Empirical Work on Aspects of Corrigibility that are Unintuitive to LLMs, and other Preliminary Results (Unpublished)(10:22) Ian Kahn -- $6,000(10:56) Various Essays on Obedience(11:00) Seth Herd -- $3,000(11:49) The corrigibility basin of attraction is a misleading gloss(11:54) Jeremy Gillen -- $2,000(12:32) A Structural Similarity Between Two Open Corrigibility Questions and Why Should Corrigible Agents Favor the Present?(12:40) Ben Saudek -- $2,000 The original text contained 3 footnotes which were omitted from this narration. ---
First published:
September 30th, 2026
Source:
https://www.lesswrong.com/posts/3uJqhrC2idf4eNj5h/corrigibility-prizes-for-existing-work
---
Narrated by TYPE III AUDIO.
---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Tristian Buckmaster recently gave an interview with Brady Haran of Numberphile discussing what happened in the Navier-Stokes drama and some context about his research. I'm posting the transcript below for people who prefer reading to watching it. It was lightly edited for clarity with Sonnet 5.5. My own view remains that it's pretty bad form for OA and other labs to race to scoop the results of researchers, and this sets a bad precedent for the future. I think they mislead Buckmaster about their swarm setup and the scale of their effort, and they could have done a better job citing previous work. However, it seems unlikely, but not implausible, that the OA access to the codex session was a major contributor to their proof. Brady: Have you got any more questions, or are you just like, "Go on, do it"? [laughter] Tristan: Yeah, just do it. Brady: You're laughing and smiling, which brings me to my first question: how are you feeling at the moment? Tristan: Better than a few weeks ago. It still hasn't calmed down, but certainly better. I have a two-month-old baby, and the week that everything happened, I probably averaged two hours' sleep [...] ---
First published:
September 30th, 2026
Source:
https://www.lesswrong.com/posts/DgnQRGeGD27unRcGv/tristan-buckmaster-numberphile-interview-transcript
---
Narrated by TYPE III AUDIO.
Summary I ran GPT-6 Sol and GPT-6.1 Sol on the task suite from Think Fast. Surprisingly, 6.1 Sol performs substantially better than 6 Sol, almost matching the performance of GPT-6 Astra. The plot below gives a quick overview of the results: Measured by mean accuracy across the 27 tasks, GPT-6.1 Sol closes 80% (95% CI: 72–87%) of the gap between GPT-6 Sol and GPT-6 Astra, and is closer to Astra than to GPT-6 Sol on 24 of 27 tasks. The likely reason behind this gap is that, like GPT-6 Astra and unlike GPT-6 Sol, GPT-6.1 Sol is a looped transformer. After providing a more detailed overview of the benchmark scores, I'll briefly discuss the evidence for this, as well as the implications. Detailed results Similarly to Astra, GPT-6.1 Sol saturates many of the benchmarks in the task suite, rendering the time horizon estimates highly uncertain. For this reason, I mainly focus on per-benchmark performance, which already provides a sufficient demonstration of the gap between 6 Sol and 6.1 Sol on its own. The time horizon estimates were 4.0 minutes for GPT-6 Sol (bootstrap median 3.8 min, 95% CI [1.2 min, 20 min]) and 35 minutes for GPT-6.1 Sol (bootstrap [...] ---Outline:(00:15) Summary(01:27) Detailed results(04:26) What caused the jump?(05:41) Appendix: Full per-task results(05:46) GPT-6 Sol: per-benchmark 50% no-CoT time horizons(06:12) GPT-6.1 Sol: per-benchmark 50% no-CoT time horizons(06:37) Appendix: Logistic fits The original text contained 4 footnotes which were omitted from this narration. ---
First published:
September 30th, 2026
Source:
https://www.lesswrong.com/posts/LqSZZAriGqgsGDQe3/gpt-6-1-sol-nearly-matches-the-no-cot-performance-of-gpt-6
---
Narrated by TYPE III AUDIO.
---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Subtitle: And maybe second best is AI safety? Further reading: So many things, but: Gradual Disempowerment, The Normalization of Deviance in AI Development, Let's Think About Slowing Down AI, Doom as a bad method, not a utopia tradeoff, Teleoperated Humans Thank you to JennaS for extensive edits and long-term discussion. I’ve been trying to get more writing out at 90% of the quality I’d like it to be at, instead of spending a bunch more time trying to wring out the last 10%, so a lot of points that could themselves be full articles are underdeveloped. Insofar as you find this post outlines a plausible or probable model of reality, or one worth criticizing centrally, let's work on developing it. Is Anthropic accelerating capabilities more than it was a year ago? At its founding? Is OpenAI accelerating capabilities more than it was a year ago? At its founding? Is GDM "laser-focused at the frontier" in pursuing recursive self-improvement? What? Why? Have they solved alignment without telling us? Why does Thomas Kwa, formerly at METR and now working on "measuring and modeling RSI" at OpenAI, worry about working at OpenAI potentially driving him (metaphorically?) insane? How is it possible [...] ---Outline:(06:34) Political Misalignment(09:06) Cultural Misalignment(15:20) Economic Misalignment(23:12) What about AI safety researchers?(25:19) Takeaways The original text contained 10 footnotes which were omitted from this narration. ---
First published:
September 29th, 2026
Source:
https://www.lesswrong.com/posts/jbttuCF4wFZmXakcj/the-world-s-best-gradual-disempowerment-model-organism
---
Narrated by TYPE III AUDIO.



