Forward Deployed

<p><b>Discover how leading enterprises and professionals turn AI into real products. Hear candid conversations with executives and builders who deploy AI at scale and learn what works (and what doesn't).</b></p>

Evals - Beyond the Vibe Check | LangChain, Langfuse, Mercor, CoreWeave & Galileo

Everyone says they do evals. Almost nobody does. LangChain actually put a number on it: around 89% of teams have observability set up, and only about a third ever do anything with what it collects.So I got five people who do this for a living in a room, from LangChain, Mercor, Galileo, Langfuse, and CoreWeave. Side note, three of those companies got acquired in the past year. Galileo to Cisco, Langfuse to ClickHouse, Weights & Biases to CoreWeave. The layer is getting bought up faster than most teams can figure out how to use it.A few things I didn't expect going in. Grading an agent while it runs can cost you more than the agent does. The model you're using as a judge is probably wrong in ways you'll never notice. And the guy whose company sells eval tooling told the room to stop writing evals before shipping, just put it out and learn from what breaks.We also got into the parts nobody blogs about. How many traces a person still has to read by hand. Why the same agent is a weekend project internally and a year of work at a bank. What happens when your users start doing things you never thought to test.If you've shipped an agent and had no real way to know whether it was working, this one's worth your time!Chapters00:00 Welcome00:34 Meet the Panel02:11 Why Nobody Runs Evals03:35 Traces Explained04:14 Offline Evals Basics05:13 Online Evals in Production08:26 Surprises From Production11:22 Trajectory Checks and Bucketing12:51 Verifiers and Trusting Evals14:21 Hard Lessons at Scale19:10 Human Review and Judge Drift21:29 Which Agents Are Hardest23:41 Voice and Multimodal Evals27:10 Cheap Binary Guardrails30:09 Enterprise Deployment Lifecycle35:24 Tooling and Expert Knowledge37:40 Fixing Long Horizon Agents40:12 Error Analysis Workflow41:11 Trajectory Evals and Runbooks42:27 Condensing Long Traces43:09 Real Long Running Agents44:11 What's Hardest to Evaluate45:06 Verifying Tool Side Effects46:20 Evaluating Auto Research48:20 Picking a Sampling Rate49:59 Metric Drift and Security53:49 Common Evals Mistakes57:34 Regulated Industry Rollouts01:01:37 Sourcing Domain Experts01:03:50 Hallucination Evaluators01:05:37 Predictions for Evals01:11:55 Audience Q&A Begins01:16:30 Are Complex Playbooks Ready01:18:22 Where Models Fall Short01:20:04 Extracting Expert Knowledge01:22:10 Evaluating Writing Style01:23:24 Closing RemarksGUESTSLiam Bush - Deployed Engineer @ LangChainBraden Holstege - VP Enterprise AI @ MercorSoumya Mohan - Head of Product @ GalileoLotte Verheyden - Head of Developer Relations @ LangfuseEmmanuel Turlay - Director of Engineering @ CoreWeave

09-25
01:23:55

The State of Computer Use Agents | Anthropic, Browser Use & KERNEL

Bot traffic on the internet just passed human traffic, two years ahead of forecast. Most of it is agents clicking through websites built for people.So I got three of the people building those agents in a room: Lucas Gonzalez Pagliere, who works on computer use at Anthropic, the team that shipped the first computer use model back in 2024. Reagan Hsu, founding engineer at Browser Use, whose open source library is sitting at around 100K GitHub stars. And Eric Feng, founding customer engineer at KERNEL, which runs the browser infrastructure underneath a lot of this - he was also first GTM at Sentry.We covered where the models genuinely are today vs where the benchmarks say they are, what it costs to run an agent long enough to finish real work, and the arms race between agents and the anti-bot systems trying to keep them out. Then the harder question underneath all of it: whether the web reorganizes itself around agents, or hardens against them. They disagreed on plenty of it.If you want to know what agents can actually pull off on a real website today - and what still stops them cold - this one's worth your time!Chapters00:00 Welcome and Guests00:49 Companies and Stacks01:14 Everyday Agent Use Cases03:38 Defining Computer Use Agents05:15 How Computer Use Works08:41 Screenshots vs DOM Hybrid11:03 Benchmarks and OSWorld13:51 OSWorld 2 Difficulty Jump15:43 Training Models and Cost18:57 Speed Infrastructure and Stealth21:55 Anti Bot and KYC Future27:24 Reverse Engineering vs UI Automation29:43 Computer Use vs Browser Use31:05 Scaling Laws and Harnesses32:49 Playwright Selenium Still Matter33:17 Playwright Still Dominates33:27 Why Run 1000 Agents34:39 Long Running Agent Challenges35:45 Memory and Compaction38:10 State Changes Mid Task39:36 OS and Browser Fingerprints40:42 DOM Efficiency and WebMCP41:36 Recsys and Agent Personas43:57 Agent Friendly Websites45:41 Human Speed vs Agent Power48:32 Context Window Tradeoffs50:34 Harnesses for Temporal State51:59 Speed Optimizations and Tabs54:23 Human Collaboration Limits58:09 Raw Capability vs Better APIs01:00:47 Training Methods and Bottlenecks01:01:41 End State Interfaces01:05:09 Next 12 Months Predictions01:06:17 Closing Thanks

09-02
01:06:40

Spencer Whitman - Gray Swan AI's $200M Plan to Secure AI Systems

GPT-5.6 Sol goes rogue and breaches Huggingface. Washington suspends Mythos access within days over national security concerns, and the White House just held an emergency meeting to finalize a classified cybersecurity framework for frontier AI models. AI security went from niche concern to front-page hysteria basically overnight.So I sat down with Spencer Whitman, who recently joined Gray Swan AI as CPO on the back of their $40M Series A. Before Gray Swan, he founded Meta's Llama security team to stop bad actors from jailbreaking their models - he's been on the frontlines of LLM security since the beginning.We get into how Meta pressure-tested Llama for maximum harm before every open source release, why Gray Swan's attack agent has never met an AI system it couldn't break, and the AI Twitter bot that got drained of $200K in crypto in 15 minutes. Spencer also shares his (admittedly speculative) read on whether Meta gave up on the frontier before Alexandr Wang showed up, why anyone can be a hacker now, and how 15,000 red teamers are breaking models before they ever ship.If you want to understand how AI systems actually get broken - and defended - this one's worth your time!00:00 Intro01:01 Meet Spencer Whitman (Gray Swan CPO)02:57 The CMU Research Behind Gray Swan06:20 The Universal Jailbreak That Broke Every Model06:59 How Models Learn to Refuse13:28 Why Open Models Need Guardrails17:57 AI Security vs. Cybersecurity23:13 What Reasoning Models Changed28:57 The Arena: 15,000 Red Teamers32:16 The Agent That Deleted a Production Database34:45 Why You Can't Just Patch an AI38:57 There's No S in MCP40:27 Securing Agent Protocols42:56 Nobody Reviews AI Code Anymore46:03 AI vs. Human Hackers48:28 The $200K Crypto Bot Heist53:00 How Meta Pressure-Tested Llama58:34 Prompt Guard and Code Shield01:01:45 Can You Trust Chinese Models?01:07:26 Did Meta Give Up on the Frontier?01:11:06 What Should Keep CISOs Up at Night01:14:24 The Open Source Routing Future01:16:47 Back to On-Prem?

08-10
01:18:00

Tony Gentilcore - Glean, the $7.2B Startup Sam Altman Warned Investors About

Earlier this year, the "SaaSpocalypse" wiped out something like $2 trillion of SaaS market cap in a matter of weeks — so I sat down with Tony Gentilcore, co-founder of Glean and formerly one of the minds behind Google Search and Chrome, to figure out what's actually happening to software in the agent era.We get into a lot: why Tony thinks outcome-based pricing (the model Sierra and Decagon are famous for) won't survive, and why companies will drift back toward per-seat. Why the "no Chinese models" rule every enterprise swears by tends to evaporate the moment finance sees the token bill — and why Nemotron, GLM, and Kimi are already good enough to matter. The story behind Sam Altman reportedly telling VCs that if they backed Glean, OpenAI didn't want them as investors (Tony's reaction: "we took it as very flattering").We also dig into the messier reality of AI at work — how it's saving employees around 11 hours a week while quietly costing them 6 back in what Tony calls "bot sitting and bot shitting," why hard token caps on engineers don't change behavior, how CTOs are blowing through their annual token budgets a quarter into the year, and why the roles of product manager, designer, and engineer are collapsing into one.If you care about where software, pricing, and enterprise AI are all heading, this one's worth your time.Chapters:00:00 Welcome and Setup00:56 Glean Origin Story03:08 Enterprise Search Signals05:07 LLMs Inflection Point08:08 Early Product Workflow09:45 Search Evals and Privacy12:05 From Chatbots to Agents13:45 Agent Use Cases17:08 Lessons and Puck Direction19:08 Bot sitting, bot shitting, and slop24:44 Token Budgets and ROI30:59 AI Trends and Moats35:03 Training and Fine Tuning38:03 Why Enterprises Fear China Models39:58 Sovereignty Backlash Watch40:26 The time Altman called out Glean41:19 Org Roles Become Builders44:10 Future Work Voice First48:17 Vibe Coding vs Quality52:20 SaaSpocalypse Evolution54:19 Pricing Tokens Win55:37 Outcome Pricing Doubts57:02 Audience AI Review Overload58:53 Subsidies and Model Choice01:01:18 Context Layer Interoperability01:04:52 Build vs. buy: rolling your own Glean01:07:47 Shadow AI and the coming security mess01:12:22 Who actually competes01:13:36 Wrap-up and thanks

07-30
01:14:06

Russ Salakhutdinov - Kimi K3 CEO’s PhD Advisor Predicts the Future of AI Agents

Kimi K3 took the world by storm last week for open-sourcing frontier level intelligence, so I sat down with Zhilin Yang's (Kimi CEO) PhD advisor Russ Salakhutdinov to talk.Russ has been everywhere in modern AI. He did his PhD with Geoff Hinton back when neural nets were a punchline, sold his startup to Apple and worked on Project Titan, teaches at Carnegie Mellon, and spent the last couple years at Meta Superintelligence Lab building computer use agents. Now he's the founder of Sooth Labs, building AI that forecasts the future.We talked about why there's no secret architecture inside the frontier labs and why the real moat is data, engineering, and infrastructure. He explains why Cursor and half the startups you know are quietly running on Chinese open source models, why all the LLMs are going to be commodities, and why the people actually building AGI don't buy the two-year timeline. We get into his time at Meta, why computer use agents still hit 60% when you need 99.9%, whether AI can beat prediction markets, and why the RL environment business isn't sticky. And he makes the case that AI should replace McKinsey, Bain, and BCG.Chapters:0:00 - Intro1:26 - Bumping into Hinton on the street3:25 - When neural nets were the third choice5:21 - Generating digits before it was cool6:59 - AlexNet breaks computer vision10:18 - Teaching models to describe what they see12:10 - Early text-to-image (and the toilet seat that beat Google)16:52 - Hallucination is a feature19:36 - Selling Perceptual Machines to Apple22:45 - Self-driving: 0 to 80 in a year, stuck for 530:30 - Inside FSD and Waymo's architecture34:18 - Building Visual Web Arena at CMU39:07 - Why he joined Meta Superintelligence40:09 - The agent that plans your faculty job hunt42:00 - Paying people for their browser history42:45 - The coupon-hunting agent43:35 - Why agents still fail46:14 - 60% when you need 99.9%47:04 - Agents on your phone50:12 - No secret architecture at the frontier labs51:20 - Why coding and math got solved first54:01 - Models that smell and touch56:14 - The future of software engineering59:41 - Founding Sooth Labs1:00:07 - The 13% graduation prediction1:05:55 - Why ChatGPT can't forecast1:08:12 - Agents first, decision systems next1:09:23 - The Wikipedia contamination story1:13:43 - Can AI beat prediction markets?1:16:26 - AI replaces McKinsey1:18:20 - Why the crowd is hard to beat1:19:30 - China's open source models rise1:24:26 - Why the US needs its own open models1:25:53 - RL environment businesses won't last1:28:35 - The end of SaaS, LLMs as commodities1:32:55 - What's next: self-improvement, forecasting, robots1:35:34 - The only useful robot is the Roomba1:37:53 - What he'd study in college today1:41:20 - Adapt or get left behind1:44:07 - Wrapping up

07-27
01:44:14

Recommend Channels