“Weight-sparse transformers have interpretable circuits” by leogao

Update: 2025-11-13

Description

TL;DR: We develop a novel method for finding interpretable circuits in Transformers, by training them to have sparse weights. This results in models that contain very high quality circuits: our circuits are global rather than datapoint dependent; we explain the circuit down to very granular objects, like individual neurons and attention channels, rather than entire MLP layers, attention heads, or groups of nodes; and the circuits are often simple enough to draw in their entirety on a whiteboard. The downside is that our method produces de novo sparse language models, which are extremely expensive to train and deploy, making it unlikely that we will ever be able to use this method to directly pretrain frontier models. We share preliminary results on using sparse models to explain an existing dense model, but our main theory of impact is to eventually scale our method to train a fully interpretable moderate-sized model. If we could fully interpret even (say) a GPT-3 level intelligence, it could aid dramatically in developing a theory of cognition in general.

[Blog] [Paper] [Code]

Abstract

Finding human-understandable circuits in language models is a central goal of the field of mechanistic interpretability. We train models to have more understandable [...]

---

First published:

November 13th, 2025

Source:

https://www.lesswrong.com/posts/yQMQXFAK4mfJjHBpN/weight-sparse-transformers-have-interpretable-circuits

---

Narrated by TYPE III AUDIO.

Comments

In Channel

“AI Corrigibility Debate: Max Harms vs. Jeremy Gillen” by Liron, Max Harms, Jeremy Gillen

2025-11-1402:24:59

“10” by Ben Pace

2025-11-1407:55

“Everyone has a plan until they get lied to the face” by Screwtape

2025-11-1412:49

“The rare, deadly virus lurking in the Southwest US, and the bigger picture” by eukaryote

2025-11-1431:38

“Creditworthiness should not be for sale” by habryka

2025-11-1414:53

“Types of systems that could be useful for agent foundations” by Alex_Altair

2025-11-1409:14

“The Charge of the Hobby Horse” by TsviBT

2025-11-1410:51

“Two can keep a secret if one is dead. So please share everything with at least one person.” by habryka

2025-11-1403:57

“Why Truth First?” by johnswentworth

2025-11-1411:54

“Orient Speed in the 21st Century” by Raemon

2025-11-1405:57

“Tell people as early as possible it’s not going to work out” by habryka

2025-11-1403:20

“Epistemic Spot Check: Expected Value of Donating to Alex Bores’s Congressional Campaign” by MichaelDickens

2025-11-1411:50

“(Fantasy) -> (Planning): A Core Mental Move For Agentic Humans?” by johnswentworth

2025-11-1403:59

“Weight-sparse transformers have interpretable circuits” by leogao

2025-11-1302:48

“What’s so hard about...? A question worth asking” by Ruby

2025-11-1304:48

“Paranoia rules everything around me” by habryka

2025-11-1322:33

“Favorite quotes from ‘High Output Management’” by Nina Panickssery

2025-11-1310:19

“The Pope Offers Wisdom” by Zvi

2025-11-1316:25

“Introducing faruvc.org” by jefftk

2025-11-1201:31

“Please, Don’t Roll Your Own Metaethics” by Wei Dai

2025-11-1204:12

00:00

“Weight-sparse transformers have interpretable circuits” by leogao

#box-pro-ellipsis-176315975987219{-webkit-line-clamp:2;}“Weight-sparse transformers have interpretable circuits” by leogao

“Weight-sparse transformers have interpretable circuits” by leogao

“Weight-sparse transformers have interpretable circuits” by leogao