Mixture-of-Experts and Trends in Large-Scale Language Modeling with Irwan Bello - #569

Update: 2022-04-25

Description

Today we’re joined by Irwan Bello, formerly a research scientist at Google Brain, and now on the founding team at a stealth AI startup. We begin our conversation with an exploration of Irwan’s recent paper, Designing Effective Sparse Expert Models, which acts as a design guide for building sparse large language model architectures. We discuss mixture of experts as a technique, the scalability of this method, and it's applicability beyond NLP tasks the data sets this experiment was benchmarked against. We also explore Irwan’s interest in the research areas of alignment and retrieval, talking through interesting lines of work for each area including instruction tuning and direct alignment.

The complete show notes for this episode can be found at twimlai.com/go/569

Comments

In Channel

High-Efficiency Diffusion Models for On-Device Image Generation and Editing with Hung Bui - #753

2025-10-2852:23

Vibe Coding's Uncanny Valley with Alexandre Pesant - #752

2025-10-2201:12:06

Dataflow Computing for AI Inference with Kunle Olukotun - #751

2025-10-1456:37

Recurrence and Attention for Long-Context Transformers with Jacob Buckman - #750

2025-10-0756:53

The Decentralized Future of Private AI with Illia Polosukhin - #749

2025-09-3058:06

Inside Nano Banana 🍌 and the Future of Vision-Language Models with Oliver Wang - #748

2025-09-2301:03:39

Is It Time to Rethink LLM Pre-Training? with Aditi Raghunathan - #747

2025-09-1658:29

Building an Immune System for AI Generated Software with Animesh Koratana - #746

2025-09-0901:04:41

Autoformalization and Verifiable Superintelligence with Christian Szegedy - #745

2025-09-0201:11:18

Multimodal AI Models on Apple Silicon with MLX with Prince Canuma - #744

2025-08-2601:09:50

Genie 3: A New Frontier for World Models with Jack Parker-Holder and Shlomi Fruchter - #743

2025-08-1901:00:31

Closing the Loop Between AI Training and Inference with Lin Qiao - #742

2025-08-1201:00:40

Context Engineering for Productive AI Agents with Filip Kozera - #741

2025-07-2945:31

Infrastructure Scaling and Compound AI Systems with Jared Quincy Davis - #740

2025-07-2201:12:32

Building Voice AI Agents That Don’t Suck with Kwindla Kramer - #739

2025-07-1501:12:32

Distilling Transformers and Diffusion Models for Robust Edge Use Cases with Fatih Porikli - #738

2025-07-0959:59

Building the Internet of Agents with Vijoy Pandey - #737

2025-06-2456:31

LLMs for Equities Feature Forecasting at Two Sigma with Ben Wellington - #736

2025-06-1759:01

Zero-Shot Auto-Labeling: The End of Annotation for Computer Vision with Jason Corso - #735

2025-06-1057:01

Grokking, Generalization Collapse, and the Dynamics of Training Deep Neural Networks with Charles Martin - #734

2025-06-0501:25:37

00:00

Mixture-of-Experts and Trends in Large-Scale Language Modeling with Irwan Bello - #569

#box-pro-ellipsis-176226024559880{-webkit-line-clamp:2;}Mixture-of-Experts and Trends in Large-Scale Language Modeling with Irwan Bello - #569

Mixture-of-Experts and Trends in Large-Scale Language Modeling with Irwan Bello - #569

Sam Charrington

Mixture-of-Experts and Trends in Large-Scale Language Modeling with Irwan Bello - #569