We explore the resurgence of self-hosted LLMs and why they're worth considering despite the convenience of frontier models. From data privacy concerns to cost management and consistent performance, there are solid reasons to run your own models. We break down the hardware requirements (spoiler: it's all about memory bandwidth), compare different runtime environments from beginner-friendly LM Studio to production-ready VLLM, and discuss practical model choices. Luca shares his experience running dual RTX 3090s in his basement, while Ryan questions whether his 2004 Athlon is still viable (it might be). We also touch on the reality that local models lag frontier models by about six months—which means they're roughly where Claude Opus 4A was in February, and that's perfectly usable for many tasks.Key Topics:[02:30] Why self-host? Data privacy, cost control, and avoiding unpredictable changes from hosted providers[08:45] Memory bandwidth is the real bottleneck—why GPUs still win and what that means for your hardware choices[15:20] Hardware options: Mac Minis vs. discrete GPUs, and why second-hand RTX 3090s offer the best price-performance[22:10] Available models: Quen, Gemini, DeepSeek, and Mistral—what works well locally and the six-month lag behind frontier models[28:40] Runtime environments compared: llama.cpp for flexibility, LM Studio for ease of use, Olama for Docker-like experience, and VLLM for production[35:15] Practical considerations: context windows, model size vs. available VRAM, and why speed matters more than you think[40:30] Announcement: UddleSpec SDD framework now available, and the Agile Embedded Podcast is now Pragmatic EmbeddedNotable Quotes:"You need memory bandwidth. All of the memory. Every single token filters through all of those neural network layers—potentially 8 billion parameters per token." — Luca"The marketing behind these tools is very much geared toward 'don't worry what's happening behind the curtain.' But there are more efficient ways of solving the same problem without incurring the same cost." — Ryan"I really feel like the most rational choice is a second-hand 3090 still. They have 24 gigabytes of VRAM, they're nice and fast, and they have about half the bandwidth of a 5090—which is still pretty good and very much usable." — LucaResources Mentioned:llama.cpp - Command-line LLM runtime environment with broad hardware support and REST APILM Studio - User-friendly GUI for running local LLMs with built-in model downloader and recommendationsOlama - Docker-like interface for managing and running local LLMs, good for experimentationVLLM - Production-grade LLM runtime with excellent parallel request handling and memory efficiencyHugging Face - Repository with hundreds of thousands of models, quantizations, and variants available for downloadUddleSpec - Luca's new SDD framework addressing real-world workflow needs—now available on GitHubPragmatic Embedded Podcast - Sister podcast (formerly Agile Embedded) covering embedded systems developmentAgile Embedded Podcast Slack - Community discussion channel where you can reach Ryan and Luca, with a dedicated sub-channel for Embedded AI topics
Ryan and Luca pull back the curtain on large language models, explaining what's really happening when you chat with an AI. They break down tokenization, context windows, and the stateless nature of LLMs—revealing why these tools aren't actually thinking or remembering, just generating the most likely next token based on massive matrices of weights. The conversation covers practical implications like why long sessions deteriorate, how caching affects costs, and why that apologetic "I won't do it again" from your LLM is meaningless. They also tackle the misleading anthropomorphization in AI marketing and share frustrations with features like Claude's opaque memory system. If you've ever wondered why your LLM seems to forget instructions or why it confidently states nonsense, this episode explains the mathematical reality behind the conversational illusion.Key Topics:[02:30] Tokenization: How LLMs break language into mathematical units[08:45] LLMs as stochastic algorithms: Predicting the next most likely token[12:20] Training models: Billions of parameters and matrix multiplication[15:10] Quantization: Trading precision for memory efficiency[18:00] Context windows and token limits: Why size matters and costs money[22:15] Stateless processing: LLMs don't remember, they reprocess everything[26:40] Caching and timeouts: The hidden costs of pausing your session[30:00] Hallucinations aren't bugs—they're the default mode of operation[35:20] Attention mechanisms and why large context windows cause degradation[40:15] Claude's memory system: Well-intentioned but problematic in practice[44:30] Mixture of experts: Sparse vs. dense models and routing tokensNotable Quotes:"The LLM doesn't even know what truth is, so it can't lie to you. It's not lying to the truth either. It's bullshitting—it doesn't care which way is true or false." — Ryan"All LLMs know how to do is hallucinate. They just generate chains of tokens. If you're lucky, those chains have some connection to the real world and are actually helpful." — Luca"I wish it wasn't programmed to just lie to me. If the LLM says 'I won't do it again,' yes it will, because it has no memory of this incident. It will behave the same way tomorrow." — LucaResources Mentioned:Luca's Training and Consulting - Luca's website with links to AI and embedded systems training coursesTulip Tree Tech Emulator - Ryan's commercial emulator for embedded systems development (Raspberry Pi, STM, ESP, Microchip)Google's Attention Paper - The foundational 2013 paper introducing the attention mechanism that enabled modern LLMsOn Bullshit by Harry Frankfurt - Philosophical work distinguishing bullshitting from lying—relevant to understanding LLM outputsAgile Embedded Podcast Slack - Community discussion channel where you can reach Ryan and Luca, with a dedicated sub-channel for Embedded AI topics
We talk with Sebastian Boblest and Marco Mader Mendes from Bosch about the surprisingly challenging art of deploying AI models on resource-constrained embedded devices. They share insights from their work on ETAS Embedded AI Coder, a tool that generates optimized C code from neural network models for microcontrollers—sometimes with as little as 100 bytes of RAM available.The conversation covers practical strategies for model compression (often by factors of 100x or more), the counterintuitive benefits of float models over quantized ones on tiny devices, and why feature engineering still matters. Sebastian and Marco explain how they navigate the trade-offs between RAM, compute time, and accuracy, and why each project presents unique constraints—from confidential hardware specs to compiler quirks. They also discuss real-world applications, including Bosch's AI-powered wall scanner that uses radar and neural networks to detect cables in walls.Key Topics:[03:30] Why Bosch started exploring embedded AI on tiny hardware in 2020[06:45] Real-world application: AI-powered wall scanner using radar to detect cables[09:20] How ETAS Embedded AI Coder generates C code from neural network models[12:00] Deploying models with as few as 300 parameters and 100 bytes of RAM[16:30] Model compression strategies: squeezing networks down by 100x or more[21:15] When float models outperform quantized ones on tiny devices[28:00] Trading off RAM vs. compute time through code generation techniques[33:45] Benchmarking challenges with confidential hardware and compilers[40:20] AutoML and architecture search for constrained embedded targets[46:00] Free tool access for universities and opportunities for studentsNotable Quotes:"We go down to applications where we use 100 bytes of RAM and you can still do something useful with this on a Cortex-M0. It's very surprising how small you can get with neural networks." — Sebastian Boblest"Sometimes we have to squeeze it down not by one or two X, sometimes it's up to 100 X and more. This is really a regular task for us." — Marco Mader Mendes"One thing that's super counterintuitive for many people on these smaller devices is to go from a quantized model to a float model. It can actually help you. It's exactly the opposite of what people do on these larger systems." — Sebastian BoblestResources Mentioned:ETAS Embedded AI Coder - Code generation tool for deploying neural networks on embedded devices; free for universitiesARM CMSIS-NN - Library containing functions for neural network layers on Cortex-M devicesMLPerf Tiny Benchmarks - Benchmarking suite for tiny ML systems that Sebastian's team participated inBosch AI-powered Wall Scanner - Radar-based power tool using neural networks to detect cables in wallsAgile Embedded Podcast Slack - Community discussion channel where you can reach Ryan and Luca, with a dedicated sub-channel for Embedded AI topics
Ryan and Luca tackle the hot topic of AI-driven security research, sparked by the release (and brief containment) of Anthropic's Mythos tool. Ryan, drawing on his 20+ years in cybersecurity, delivers a surprisingly reassuring message: if you've been doing proper engineering, AI-found vulnerabilities aren't the existential threat the hype suggests.We explore how Mythos and similar tools are flooding the CVE database with thousands of new vulnerabilities, but discuss why this doesn't automatically mean more successful attacks. Ryan explains the crucial difference between finding a bug and actually exploiting it, and why embedded systems developers shouldn't panic—but should definitely have their update processes sorted. The conversation covers everything from botnet refrigerators to Ukrainian security cameras, threat modeling with AI assistance, and why AI makes such a relentlessly effective hacker (spoiler: it doesn't get bored). Bottom line: the weapons haven't changed, you just need to install the bulletproof glass you should have had all along.Key Topics:[02:30] Introducing Mythos: AI tool for finding software vulnerabilities, its May release, US containment, and recent re-release[05:15] The flood of AI-generated bug reports: distinguishing real vulnerabilities from noise, and the burden on maintainers[08:45] Mythos by the numbers: 6,000+ critical vulnerabilities found, 90% true positive rate, but does it matter for your system?[12:00] The reachability problem: having a bug vs. being able to exploit it, and why solid engineering processes matter more than bug counts[16:30] Embedded systems challenges: BSP version conflicts, regulatory approval nightmares, and the CRA/FDA compliance push[21:00] No uptick in actual attacks: why more CVEs doesn't equal more breaches, and what motivates attackers (hint: money, not bugs)[28:45] Embedded systems' blessing and curse: air-gapped devices vs. internet-connected vulnerabilities, and the botnet refrigerator story[35:20] Using AI for defense: network analysis, threat modeling, traffic pattern recognition, and why AI is the best grep you've ever seen[42:00] AI as the relentless attacker: no social norms, no contracts, just pure problem-solving—and why that's both powerful and concerning[47:30] The offensive vs. defensive mindset: why AI bridges both motivational patterns and what that means for security teamsNotable Quotes:"Just because there's more CVEs doesn't really change the defensive posturing that you need to do. If you have a solid engineering process to provide updates to your system and test and verify those updates actually work, it doesn't matter how many bugs you find. You can find two. You can find 6,000." — Ryan Torvik"AI doesn't have social norms. AI will just go and do things. And yes, there are guardrails that they're trying to put on, but like, I don't know if you've seen AI just ignore parts of your prompt before." — Ryan Torvik"There's not a death star. This is not an existential crisis. This is not something new. It changes the game but not in a way that we can't handle. The weapons haven't changed, you just need to install the bulletproof glass." — Ryan TorvikResources Mentioned:Anthropic's Mythos - AI tool for automated vulnerability discovery in software, released in May 2024, briefly restricted by US authorities, then re-releasedCVE Database - Common Vulnerabilities and Exposures database, currently experiencing significant growth due to AI-assisted vulnerability discoveryAgile Embedded Podcast Slack - Community Slack channel where Ryan participates, sister podcast to Embedded AITulip Tree Tech - Ryan's company focused on improving embedded development processes, including emulator-in-the-loop solutionsluca.engineer - Luca's website with links to training courses, LinkedIn, and other professional activities
We tackle the elephant in the room: will AI take our jobs? Spoiler alert - probably not, but your job will definitely change shape. Ryan and Luca dig into what software development actually is (hint: it's not just mashing keyboards), why embedded systems might be particularly safe from AI disruption, and what we can learn from the Luddites and steam engines.Drawing on Luca's experience with DevOps transformations and Ryan's work running a company, we explore why engineers keep solving the wrong problems, what Black & Decker actually sells (it's not drill bits), and why vibe-coding your email server is a terrible idea. The real question isn't whether AI will replace you - it's whether you understand what your actual job is in the first place. Plus: Pokemon evolution as career advice, and why AI is the new Agile fairy dust.Key Topics:[02:30] The real job vs. the mechanical task - why writing code isn't actually your job[08:45] Why embedded systems are particularly safe from AI disruption - complexity, hardware interaction, and undocumented quirks[15:20] The Black & Decker lesson: selling holes in walls, not drill bits - understanding what customers actually want[22:10] Historical parallels: steam engines, DevOps, and why demand grows faster than automation[28:40] The reality check: most developers aren't using AI systematically yet - you're not behind the curve[32:15] AI as the new Agile fairy dust - why magic solutions never work without process and understandingNotable Quotes:"Your job is not writing code. Your job is making product. If you think about the Luddites, their job was not operating a loom. Their job was making clothes." — Ryan Torvik"Writing code is really the smallest part of software development. Most of it is sitting in front of the screen, looking up and to the left, and figuring out what code to write." — Luca Ingianni"The CEO of Black & Decker was once quoted as saying: We are not in the business of selling drill bits. We are in the business of selling holes in the wall. If we had laser cannons that made holes in walls, people would buy the laser cannons instead." — Luca IngianniResources Mentioned:Agile Embedded Podcast Slack - Community discussion channel where you can reach Ryan and Luca, with a dedicated sub-channel for Embedded AI topicsLuca's Website - Contact Luca for consulting on using AI effectively in embedded systems contexts - multiple contact options available