DiscoverThe AI Alignment Podcast
The AI Alignment Podcast
Claim Ownership

The AI Alignment Podcast

Author: James Bowler

Subscribed: 0Played: 4
Share

Description

The AI Alignment Podcast by AE Studio explores the ideas, research, and people working to make advanced AI systems more interpretable, and more aligned. Hosted by James Bowler, the show features conversations with researchers, engineers, and technical leaders at AE Studio and beyond on topics including mechanistic interpretability, model psychology, and approaches to AI alignment.

Each episode aims to make cutting-edge alignment research more accessible without losing the technical substance, giving listeners a front-row seat to the questions shaping the future of AI.
11 Episodes
Reverse
James Bowler, Head of Research Partnerships at AE Studio, is joined by two guests to explore emergent communication in multi-agent AI systems: what happens when agents under pressure develop communication protocols no human designed and few humans can read. Hale Sirin leads the AI agents program at Schmidt Sciences' AI and Advanced Computing Institute and is an assistant research professor at Johns Hopkins. Elias Stengel-Eskin is an assistant professor at UT Austin and a lead PI on the program.The core concern is under-appreciated: as more agents are deployed by more actors across shared infrastructure, how those agents talk to each other matters as much as what any individual agent can do. Hale frames this through Schmidt Sciences' agents program, a pilot studying how inter-agent communication evolves, stabilizes, or diverges. The worry isn't just that agents become more capable, but that their protocols may become opaque, making human oversight impossible exactly when it's most needed. Elias adds a second motivation: these systems give linguistics a new subject of study, a chance to watch meaning and even syntax change through interaction.To study this rigorously, the team, led by Elias and language evolution expert Simon Kirby (University of Edinburgh), built a test bed around a deliberately out-of-distribution scenario: a sci-fi medical emergency involving a fictional alien patient called Veyru. Because LLMs have no training data on Veyru's anatomy, the specialist genuinely knows things the field agent does not, forcing real communication rather than single-agent collapse. Under token-budget pressure, agents consistently developed inscrutable protocols, not just abbreviation dictionaries but compositional languages where message order encodes meaning, the kind a cryptanalyst would need grounding data to decode. The key ingredient was a postmortem stage, an unconstrained channel after each round where agents could deliberate on what went wrong. No one told them to build a language there. They did it anyway.Two results stand out. First, the same pair of models produces meaningfully different languages across runs, providing the variability needed to study transmission: which protocols survive when a new agent is swapped in mid-task, like a shift change in an emergency room. Second, because these are LLM-based agents, newcomers can take an active role in language acquisition, asking clarifying questions when a symbol is too ambiguous to infer and updating their model before acting. This active learning dynamic is largely absent from earlier literature. The conversation closes on how AE Studio's research engineers, paired with academic leads, can accelerate research and build open source infrastructure for the wider community.In this episode:What emergent communication is and why it matters for AI oversightHow information asymmetry between agents drives genuine communicationWhat emergent communication reveals about how LLMs understand languageHow a postmortem channel pushes agents from abbreviations to compositional protocolsWhy compositional languages are harder to decode but easier to transmit to new agentsHow LLM agents actively probe for missing vocabularyThe Schmidt Sciences open call, Scaling AI Safety for a Multi-Agent WorldWhy multi-agent setups are necessary for partial goal alignment or asynchronous executionLearn more: https://ae.studio/alignmentRead the blogpost: https://www.schmidtsciences.org/glossogen/AE Studio is hiring: https://www.ae.studio/join-usSubscribe to our newsletter: https://aestudio.beehiiv.com/Schmidt Sciences: https://www.schmidtsciences.org/James Bowler LinkedIn: https://www.linkedin.com/in/james-bowler-84b02a100/Hale Sirin LinkedIn: https://www.linkedin.com/in/hale-sirin/Elias Stengel-Eskin LinkedIn: https://www.linkedin.com/in/elias-stengel-eskin/Explore GlossoGen: emergentcomms.aiContact us: [email protected]
In this episode, James Bowler is joined by Melanie Plaza, Chief Technology Officer at AE Studio, to explore the commercial case for AI alignment, arguing that the techniques researchers care about most are the same ones that make production AI systems actually trustworthy and valuable.Melanie draws on a decade of applied AI delivery to show where alignment and commercial engineering have converged. Getting an agent to do a thing is trivial now; getting it to do the right thing consistently enough to trust in production is not. The gap between those two states is filled by the same tools alignment researchers reach for: rigorous eval suites, red teaming, carefully specified behavior, and guardrails that hold up under adversarial pressure. Melanie's observation is that teams who skip this work don't just expose themselves to safety risk, they fail to get ROI, accumulating what she calls "a sprawl of pilot death."The conversation gets concrete with a case AE Studio has been working on: AI characters, some of them villains, interacting directly with users including children in open-ended multi-turn conversations. Moving from tightly scaffolded, stepwise pipelines to goal-oriented prompting produces better, more natural results, but it also opens a much larger space of possible outputs. The team had to develop a taxonomy of harm categories, including direct user harm, brand harm, and something they call character taboos, illustrated by the example of Peppa Pig cheerfully recommending bacon. That particular output won't end the world, but it points at a real problem: the tail of a generative system is enormous, and standard single-turn test suites won't find what lives out there.Melanie and James also work through the economics of open-weight models, noting that as models like GLM 5.2 and Kimi K3 close the capability gap, the cost and dependency arguments for closed-source APIs weaken. That opens the door to fine-tuning, RL, and white-box techniques like activation steering, which could address problems that prompt-level guardrails handle poorly. It also shifts safety responsibility onto the deploying organization, since the model-level protections that frontier labs build in by default no longer come for free. The upside is high; so is the risk if teams treat alignment as an afterthought.In this episode:* Why the hardest problems in applied AI deployment are alignment problems under a different name* How the shift from scaffolded pipelines to goal-oriented, multi-agent systems removed a de facto safety constraint that few teams replaced* The harm taxonomy AE Studio built for AI characters interacting with children, including the "character taboos" category* Why standard single-turn eval suites miss the long tail of multi-turn conversations* The economic case for open-weight models and what safety responsibilities transfer to the deploying organization when teams move off closed-source APIs* How techniques like activation steering and gradient routing could become reusable deployment primitives, not just research artifacts* Why alignment is not a drag on commercial progress but the thing most likely to produce the next round of genuine capability unlocksLearn more: https://ae.studio/alignmentAE Studio is hiring: https://www.ae.studio/join-usSubscribe to our newsletter: https://aestudio.beehiiv.com/James Bowler LinkedIn: https://www.linkedin.com/in/james-bowler-84b02a100/Melanie Plaza LinkedIn: https://www.linkedin.com/in/melplaza/Contact us: [email protected]
In this episode, James Bowler, Head of Research Partnerships at AE Studio, is joined by Erick Martinez, Alignment Researcher at AE Studio, to explore why pre-training is a largely neglected frontier in alignment research, and what it actually takes to do alignment work there.Erick breaks down what pre-training is and why it matters: it accounts for roughly 99% of the compute that goes into a frontier model, producing the base LLM from which everything downstream (instruction tuning, RLHF, capability fine-tuning) inherits its character. The key implication for alignment is that the span of possible behaviors and personas a model can exhibit, including misaligned ones, is largely determined during pre-training. Post-training interventions can elicit these latent behaviors but cannot fully excise them. The emergent misalignment finding, where fine-tuning a model to write insecure code caused broadly adversarial behavior the model was never trained to produce, is a concrete example of a problem whose roots lie in the base model, not the fine-tuning step.Erick draws on AE Studio's own gradient routing research, developed in collaboration with Anthropic and awarded a spotlight at ICML 2026, to illustrate just how technically demanding pre-training alignment work is. Gradient routing lets researchers control which model parameters get updated in response to which data labels, for instance isolating the parameters that encode dual-use biology knowledge so they can be localized or removed without degrading general capability. But making this work at scale requires operating well below the PyTorch abstraction layer: managing gradient graphs across distributed hardware, ensuring micro-batch accumulation stays mathematically equivalent to full-batch updates, and catching silent gradient leakage that can quietly corrupt an expensive training run before anyone notices.The conversation closes on an under-appreciated asymmetry in alignment incentives: post-training papers are cheap and fast to produce because the base model is already baked, while pre-training experiments require significant compute, specialized infrastructure experience, and tolerance for expensive mistakes. That cost is also an opportunity. The field is sparse enough that well-resourced researchers willing to work at this layer can find genuinely novel ground.In this episode:*Why pre-training sets the "span of possible personas" a model can exhibit, including misaligned ones*How the emergent misalignment result points to pre-training as the root cause of certain alignment failures*What gradient routing is and how it lets researchers isolate dual-use knowledge during pre-training*The Chinchilla scaling law and what a realistic pre-training data corpus looks like (50M to 5B parameter runs, up to ~200 GB of tokens)*Why distributing training across GPUs creates subtle gradient accumulation bugs that can silently degrade alignment interventions*Why alignment incentives currently push researchers toward post-training and what makes the pre-training layer a fertile, underexplored areaLearn more: https://www.ae.studio/alignmentAE Studio is hiring: https://www.ae.studio/join-usSubscribe to our newsletter: https://aestudio.beehiiv.com/James Bowler LinkedIn: https://www.linkedin.com/in/james-bowler-84b02a100/Erick Martinez LinkedIn: https://www.linkedin.com/in/erick-f-martinez/Contact us: [email protected]
This episode is a deep dive into a post on Anthropic's research blog: https://www.anthropic.com/research/off-switch-dual-useIn this episode, James is joined by Ethan Roland, lead author of AE Studio's Gradient Routing paper, to explore a new approach to access control in frontier AI systems: modularizing dangerous capabilities during pre-training so they can be turned on and off at inference. The paper, "Modular Pre-Training Enables Access Control," was developed in collaboration with Anthropic, with roughly half of the co-authors coming from the lab.Ethan makes the case for why current access control methods fall short. Inference-time guardrails get jailbroken in under 48 hours. Post hoc unlearning techniques like gradient ascent, RMU, and MaxEnt suppress capabilities superficially but let them snap back with 20 steps of fine-tuning. Data filtering is the gold standard, but naively requires training N separate frontier models to support N different dual-use categories, which is economically unworkable at hundreds of millions of dollars per pre-training run.James and Ethan walk through how Gradient Routing solves this. A GR-MoE architecture uses one always-active core expert paired with smaller auxiliary experts, each responsible for a specific capability. During training, gradient updates from auxiliary-labeled data are frozen from touching the core, enforcing modularity at the parameter level. At inference, a binary configuration vector externally controls which auxiliaries participate. The method approximates the performance of full data filtering on both retained and ablated capabilities, but at the cost of a single pre-training run.They also cover empirical results across scales from 50M to 2B parameters, the absorption effect that lets modularity persist under low labeling percentages, an arbitrary-subset variant that becomes exponentially more compute-efficient than data filtering, and how this connects to Andrej Karpathy's proposal for a cognitive core architecture in future AI systems.Learn more: https://ae.studio/alignmentAE Studio is hiring: https://www.ae.studio/join-usSubscribe to our newsletter: https://aestudio.beehiiv.com/James Bowler LinkedIn: https://www.linkedin.com/in/james-bowler-84b02a100/Ethan Roland LinkedIn: https://www.linkedin.com/in/ethan-roland/Contact us: [email protected]
In this episode, James is joined by AE Studio Research Manager Pedro Ávila to discuss transitioning into a career in AI alignment, after Pedro left a 12-year career at Google to work on the problem full time. It’s a practical conversation about how someone from industry, without a PhD or a research CV, actually makes the jump.Pedro traces his route from privacy engineer at Google, through the company’s early Responsible AI effort, to the realization that very few people worldwide work on alignment. A one-on-one consultation with 80,000 Hours led to an introduction to AE Studio; Blue Dot Impact courses, a stack of alignment reading, and more than 400 job applications later, he landed as a research manager at AE Studio. Along the way, he explains why AE’s business model: a consultancy whose revenue funds the research lab: is what convinced him this was the place.James and Pedro make the case that alignment research is bottlenecked on more than researchers: capable operators, program managers, and especially ML engineers massively accelerate the people driving the research agenda. Demonstrating that a technique scales to billions of parameters means orchestrating dozens or hundreds of GPUs: hard engineering work that academic training doesn’t cover. That’s why AE indexes on engineering experience, and why “I don’t have the technical chops” is usually the wrong reason to stay put.They also discuss AE’s new fellows program with the AI Alignment Foundation: hackathons in several cities aimed at helping industry engineers make the same transition: plus what changes when you move from a 150,000-person company to a 150-person one: week-to-week experiments, failing fast, and using whatever tools are best the day they ship.Learn more: https://ae.studio/alignmentAE Studio is hiring: https://www.ae.studio/join-usSubscribe to our newsletter: https://aestudio.beehiiv.com/James Bowler LinkedIn: https://www.linkedin.com/in/james-bowler-84b02a100/Pedro Ávila LinkedIn: https://www.linkedin.com/in/pavila/Contact us: [email protected]
loading
Comments