DiscoverThe Information Bottleneck
The Information Bottleneck
Claim Ownership

The Information Bottleneck

Author: Ravid Shwartz-Ziv & Allen Roush

Subscribed: 23Played: 451
Share

Description

Two AI Researchers - Ravid Shwartz Ziv, and Allen Roush, discuss the latest trends, news, and research within Generative AI, LLMs, GPUs, and Cloud Systems.

73 Episodes
Reverse
Yuandong Tian spent many years at Meta FAIR and recently left to co-found Recursive Superintelligence, a company building AI that improves itself. In this episode, he tells us why.We start with his work on computer Go, where he built DarkForest before AlphaGo came out. A few years later came OpenGo, which played Korean professionals on a single GPU and didn't lose a game. He then tried to bring RL to real-world problems and found that the design of the action space mattered more than the algorithm.We also talk about Coconut, his paper on reasoning in latent space instead of tokens, and why frontier models still aren't trained that way.The second half is about recursive self-improvement. Yuandong wrote his last paper at Meta together with GPT-5 and says it made him 6 to 10 times faster. That convinced him his own job could be replaced within five years. We ask him where agents still fall short and whether a new architecture can really beat transformers at scale. He also gives his view on the calls to restrict self-improving AI and makes the case for open source.Recursive is hiring in San Francisco and London: [email protected]* Computer Go: DarkForest, AlphaGo and OpenGo* Gradient-free optimization* Action space design and neural architecture search* Coconut and reasoning in latent space* Understanding how neural networks learn representations* Grokking, and writing a paper with GPT-5* Recursive self-improvement and coding agents* New architectures vs. transformers* The NanoGPT speedrun* Data efficiency and robotics* Restricting self-improving AI* Open source modelsTimeline0:00 Intro0:45 From CMU to deep learning, and the AlexNet debates5:35 AlphaGo and beating Go pros on a single GPU10:27 Gradient-free optimization13:16 Diffusion models and diversity16:23 RL on real problems: why the action space matters20:34 AutoML and architecture search22:05 Coconut and reasoning in latent space26:07 Why latent reasoning hasn't caught on30:36 Adapting research to the LLM era34:04 Why he left Meta to work on recursive self-improvement37:10 Do we still need humans?39:37 Where agents fall short43:08 Can new architectures beat transformers at scale?47:29 The hardest part of research to automate50:16 What comes next for AI52:52 Should self-improving AI be restricted?54:11 Open source models56:43 Hiring at RecursiveMusic- "Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Chris Manning is a legend in NLP. He's a professor of linguistics and computer science at Stanford and ran the Stanford AI Lab. His course CS224N, Natural Language Processing with Deep Learning, is where a whole generation of researchers learned the field. If you've worked on anything in NLP over the last 25 years, you've probably used his work.We start with what linguistics gave machine learning that ML wouldn't have figured out on its own, and what's left for NLP researchers now that LLMs handle most of the classic tasks. Chris thinks the open problems have moved up the stack to pragmatics and dialogue. Models are good at using context but still sound confident when they shouldn't.Then we get into his recent work on how language models learn verb classes. Small GPT-2 models seem to form abstract categories right away instead of memorizing verbs one at a time, and Chris explains why distributed representations push them in that direction.The second half is about meaning and representation. Chris makes the case that Yann LeCun underrates the role of language in intelligence, and revisits his debate with Emily Bender over whether text alone can teach meaning. We finish with ReFT, which steers a frozen model by editing its hidden states, and whether concepts really live in linear subspaces.Timeline00:00 Intro01:05 What linguistics gave machine learning03:27 Is NLP losing its focus on language?06:29 What's left for NLP researchers in the LLM era09:37 The next frontier: pragmatics and dialogue13:01 Constructed languages and AI-to-AI communication15:35 Do language models learn categories first?23:43 Why transformers learn abstractions early26:37 Does it hold at scale?28:30 LeCun, JEPA and the role of language in intelligence34:54 Can text alone teach meaning? The octopus debate40:45 Diffusion language models vs transformers43:39 ReFT: editing representations instead of weights50:34 Weights or representations: where knowledge lives52:37 Do concepts really live in linear subspaces?TopicsNLP, computational linguistics, large language models, learning dynamics, meaning from form, world models, JEPA, diffusion LMs, ReFT, representation learning, interpretabilityMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Andrew Dai spent over a decade at Google Brain and DeepMind, where he co-wrote the 2015 paper that introduced language model pre-training followed by fine-tuning, and later co-led pre-training data for Gemini. He's now co-founder and CEO of Elorian AI, which is building models for visual reasoning.We talk about how his pre-training result started as a bug, why next-token prediction scales better than other objectives, and what went wrong for Google in the early LLM race. The second half is about vision: why today's frontier models still can't count objects in a photo, why he thinks reasoning is fundamentally visual, and how his view of world models differs from JEPA.Chapters00:00 Intro00:48 Andrew's background and the accidental discovery of pre-training07:10 Why next-token prediction scales10:09 How Google fell behind and the early days of Gemini18:28 What makes training data good28:16 Where visual understanding breaks down46:01 World models, JEPA and robotics55:27 Generation vs understanding, and what's next for ElorianTopicsPre-training and fine-tuningScaling and next-token predictionGemini and Google's AI historyData quality and synthetic dataVisual reasoning and countingWorld models and JEPAMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
Zachary Lipton is an associate professor at Carnegie Mellon University and a co-founder of Abridge, a healthcare AI company (and a jazz saxophonist!). His research spans machine learning, healthcare, and the broader impact of AI.In this episode, we discuss what AI can (and cannot yet) do in healthcare. We talk about why medicine is harder to automate than coding, AI scribes and clinical decision support, drug discovery, whether AI could eventually replace doctors, and the role of open models and intelligent routing.In the second half, we turn to the future of AI research: whether academia can still compete, what AI PhD students should work on, and how Zach thinks about automation and the future of research.Timeline00:00 — Introduction00:28 — Why healthcare is hard for AI05:00 — Zach’s path into healthcare AI22:13 — Where AI can have the biggest impact32:01 — Can AI replace doctors?41:41 — Open models and AI infrastructure56:20 — The future of AI research1:03:03 — What should AI PhD students work on?1:27:00 — Advice for researchers and foundersTopicsAI in healthcare • AI doctors • drug discovery • clinical decision support • open-source AI • AI research • academia vs. industry • future of workMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
What comes after scaling?We talk with Sara Hooker, co-founder and CEO of Adaptation Lab, about why the next generation of AI may look very different from today's static models. Sara argues that models should continuously adapt to new tasks, data, users, and environments—and that doing this efficiently will require rethinking much more than fine-tuning.We discuss continual learning, AutoScientist and automated research, why non-verifiable tasks may become the next major bottleneck, and why interfaces could be as important as the models themselves. We also get into open vs. closed models, distillation and Chinese AI labs, AI regulation and safety, cybersecurity and biorisk, AI companionship, and what may eventually come after Transformers and tokenization.TopicsContinuous learning and adaptive AIFine-tuning, memory, and AutoScientistAI agents and automated researchNon-verifiable tasks and human feedbackAdaptive interfacesOpen vs. closed models and distillationAI safety, regulation, cyber risk, and bioriskAI companionship and persuasionThe limits of TransformersMultilingual models and tokenizationChapters00:00 — Introduction02:15 — Why start another AI lab? The return of research05:46 — What continuous learning actually means12:04 — Should every company have its own adapting model?13:59 — Fine-tuning and platforms like Tinker18:04 — AutoScientist and automated optimization22:52 — Can AI really improve its own research?28:38 — The problem of non-verifiable tasks31:30 — Human feedback and the limits of exponential progress34:43 — Why the AI interface matters40:36 — Distillation, China, and open models49:05 — Open-model licensing52:19 — Will open models catch closed models?58:43 — AI regulation and compute thresholds1:03:07 — AI safety and agent failures1:10:19 — Biorisk vs. cybersecurity1:14:03 — Persuasion, AI companionship, and overlooked risks1:20:41 — Where will AI have the biggest real-world impact?1:25:41 — What is missing from current AI architectures?1:29:03 — Neurosymbolic AI1:31:30 — Multilingual models and tokenization1:34:02 — Byte-level models and alternatives to tokenization1:35:03 — ClosingMusic"Kid Kodi" - Blue Dot Sessions - via Free Music Archive - CC BY-NC 4.0
loading
Comments