Discover
Machine Learning Guide
Machine Learning Guide
Author: OCDevel
Subscribed: 19,754Played: 123,933Subscribe
Share
© OCDevel copyright 2025
Description
Machine learning audio course, teaching the fundamentals of machine learning and artificial intelligence. It covers intuition, models (shallow and deep), math, languages, frameworks, etc. Where your other ML resources provide the trees, I provide the forest. Consider MLG your syllabus, with highly-curated resources for each episode's details at ocdevel.com. Audio is a great supplement during exercise, commute, chores, etc.
60 Episodes
Reverse
The aggregate job market held, the entry-level door narrowed, and software postings sit a quarter below pre-pandemic. Why cheap implementation made specification, verification and domain scarce, how ML roles split five ways, and how to position. Links Try a walking desk - stay healthy & sharp while you learn & code More OCDevel shows - this one has siblings, each on its own subject and produced the same way What coding agents did to programming and machine learning jobs by late 2026: the labor data, the mechanism behind it, how the ML career splintered into five roles, and a concrete positioning plan. Companion to the vibe coding trio (Vibe Coding in 2026, Inside a Coding Agent, Agentic Software Engineering) and the agents pair (AI Agents in 2026, OpenClaw and the Personal Agent). Displacement vs task change Two different claims hide inside "AI is taking programming jobs": displacement (the role disappears, nobody is rehired) and task change (the role stays, the work shifts). They appear in different data. Displacement shows in unemployment and layoff reports; task change shows in what postings ask for and how teams are shaped. The aggregate evidence is mostly task change with one real pocket of displacement at the entry level, which sets up offensive advice for most listeners and defensive advice for new entrants. The evidence Aggregate. The BLS Employment Situation has unemployment at 4.1% with payrolls beating forecasts, far from the 10-20% Dario Amodei floated in the Axios "white-collar bloodbath" interview. Both he and Sam Altman have since softened the timeline; Altman said he was "delighted to be wrong" (Fortune, Time). The Yale Budget Lab tracker finds no discernible disruption. Goldman Sachs Research estimates a net drag of about 16,000 jobs a month across 800+ occupations, with a long-run baseline of 6-7% of workers displaced. Entry level. Stanford's Canaries in the Coal Mine (August 2026 paper, dashboard) puts 22-25 year olds in AI-exposed occupations 19% behind less-exposed peers, up from 15% a year earlier, driven by reduced hiring rather than separations and concentrated in automation-style exposure. The authors call these descriptive indicators, not causal estimates. Their software-developer case study finds young-developer pay grew somewhat faster than older developers' after ChatGPT, consistent with firms hiring fewer but better-paid juniors; the CPS sample is too small for a software-specific employment percentage. The EIG entry-level working paper is the main counterweight. Software demand. Indeed's software development postings index (Feb 2020 = 100) sits near 75, roughly a quarter below pre-pandemic and still drifting down, against a much smaller decline in total postings. Confounds stacked on top of AI: rate hikes, the Section 174 expensing change, and the 2021 overhire. SignalFire's State of Talent has new grads at 7% of Big Tech hires, down 25% from 2023 and over 50% from 2019, so half the collapse predates ChatGPT. Layoffs. Challenger, Gray & Christmas counts 116,175 of 529,914 announced 2026 cuts through August as AI-attributed (about 22%, already more than double all of 2025); AI led every month from March to July, then fell to 3,462 in August, while year-to-date cuts are down 41%. Grads and incumbents. The NY Fed college labor market data (2026:Q2) has computer science at 7.0% unemployment and 19.1% underemployment and computer engineering at 7.8% and 15.8%, against 5.6% and 42% for all recent graduates: worst on getting a job, among the best on getting a good one. CompTIA's tech jobs report has tech occupation unemployment at 2.8% and over 320,000 active postings asking for AI-related capabilities. The reversal wave: CNBC and Forbes on employers rehiring after AI cuts, Forrester's 55% regret figure, Robert Half's one-in-three refill figure, and the Klarna and IBM cases. Projections. BLS 2025-2035: software developers +10% from 1.72 million, data scientists +35%, computer programmers -7%. Measurement. METR's randomized trial of 16 experienced open-source developers on 246 real issues found AI made them 19% slower while they believed it sped them up 20%, so self-reported productivity is unreliable in both directions. The mechanism When implementation cost falls toward zero, value moves to specification, verification and domain knowledge. The BLS programmer-vs-developer split is that thesis in two rows. Andrew Ng's AI Rewards Generalists Who Can Build New Skills and his five-part AI Engineering Skills Map argue the bottleneck moved from how to build to what to build; David Autor calls AI a supplement to workers with judgment and domain knowledge. Juniors are hit because the traditional junior role was the commoditized part, and it was also the tuition for learning the other two skills. The ceiling on the mechanism shows in two benchmarks from the same year: OpenAI's GDPval, where the newest models win or tie against experts on most one-shot deliverables (with caveats about automated grading), against Scale's Remote Labor Index, where the best agent completed about 4% of real multi-day projects (via Carnegie). Agents produce artifacts; humans still run projects. The ML career in 2026 Data scientist demand is projected to grow three times as fast as developer demand, and Levels.fyi puts ML/AI-focused engineers in the US around $248k average total compensation. The title splintered into five roles, roughly by headcount: AI engineer (application layer: retrieval, tool use, agent loops, evals, context design); forward-deployed engineer (the Palantir-origin role the labs adopted, where domain is the constraint; see Anthropic's FDE posting); evals and AI quality (titles like Research Engineer, Model Evaluations on Anthropic's jobs board); inference, serving and platform infrastructure; and research engineer or scientist, the smallest and most competitive tier. "AI engineer" now means treating the model as a component with a failure distribution and designing the system around it. Prompt engineering as a standalone title, fine-tuning as a default move, and train-from-scratch generalist ML roles lost ground. Three camps Accelerationists (Amodei, Altman, Mustafa Suleyman): disruption within one to five years, entry-level first; strongest evidence is the benchmark curve. The aggregate prediction has failed so far and both leading voices softened it; Suleyman's 12-18 month clock has not expired. Skeptics (Yann LeCun, who left Meta to found AMI Labs on a world-model thesis; Gary Marcus; Daron Acemoglu, whose macro estimate is under 1% TFP gain over a decade): strongest evidence is the Remote Labor Index; weakest point is the cheap-but-imperfect case that reshapes jobs without replacing them. Pragmatists (Andrew Ng, Brynjolfsson, Autor): technology real, effects uneven, the question is which tasks move. Best track record so far because they predicted least. The Carnegie Endowment's three views cuts the map differently and is worth reading alongside. The Anthropic Economic Index (January, March) shows augmentation edging up on consumer chat while API and coding-agent usage stays automation-dominant. Positioning Own a domain where correct answers require knowledge not on the internet. Own verification: reading diffs fast, writing the test before the bug, building the eval harness, catching reward-hacked tests. Run agents fluently and measure yourself rather than trusting the feeling (the METR gap). Ship agentic work in public with specs, tests, evals and review trail visible, the new portfolio. New entrants: don't look like the traditional junior; compete for the well-paid junior seats that remain, at companies with real domains, in the roles that are hiring. Learning path Fundamentals first, because you can't verify what you don't understand: the Machine Learning Guide core episodes. Then the applied layer, which changes every few months: the vibe coding trio, the agents pair, the media trio. Then a domain and a project, which no course provides. Related episodes MLA 22: Vibe Coding in 2026 MLA 23: Inside a Coding Agent MLA 24: Agentic Software Engineering MLA 28: AI Agents in 2026 MLA 29: OpenClaw and the Personal Agent Every show Gnothi has produced, on AI, coding agents, video generation and agentic business, is at ocdevel.com/moremlg.
OpenClaw as the worked example of the always-on personal agent: gateway, markdown memory, heartbeats, skills, coding agents from your phone, hosted vs local models, and the 2026 security record (exposed instances, two critical CVEs, ClawHavoc) with the posture that makes it survivable. Links Try a walking desk - stay healthy & sharp while you learn & code More OCDevel shows - this one has siblings, each on its own subject and produced the same way Second and last episode of the agents pair. AI Agents in 2026 covered the theory: loops, tools, memory, protocols, SDKs, evaluation. This one takes a single category, the always-on personal agent, and its most-cloned instance, OpenClaw, all the way down to the security posture required before it touches your inbox. What a personal agent is A personal agent is the agent loop with three additions: it runs continuously on a machine you control and can wake itself; its interface is messaging (WhatsApp, Telegram, Signal, iMessage, Slack) rather than a chat tab; and it holds your files, shell, browser, calendar and, if you allow it, email. Messaging removes the gap between having a thought and delegating it, and a chat thread is a natural home for asynchronous work that reports back later. The access that makes it useful is also the entire security problem. OpenClaw today OpenClaw is an MIT-licensed, self-hosted TypeScript agent created by Peter Steinberger in late 2025. It was released as Warelay, passed through claw-themed names, hit an Anthropic trademark complaint, and settled on OpenClaw at the end of January 2026 per Wikipedia; the org remains openclaw/openclaw. Steinberger joined OpenAI in February 2026 and stewardship moved to the OpenClaw Foundation, a US 501(c)(3) chaired by Dave Morin with a small full-time staff and donors including OpenAI, GitHub, Nvidia and Microsoft; the repo's README says OpenAI is a donor, not an owner, and there is no paid tier or hosted service. The project ships near-weekly CalVer releases plus an extended-stable line, and the GitHub blog's maintainer profile called it the fastest-growing project in the site's history. OpenClaw 2.0 (v2026.8.1) landed at the end of August. Rewrites that keep the idea and shrink the surface, per OSS Insight's fork-wave analysis: nanobot (Python, small auditable core), ZeroClaw (Rust, single static binary), PicoClaw (Go, from Sipeed, embedded targets) and NanoClaw (TypeScript, container-first). Hosted versions are all third-party one-click deploys or small managed services; the foundation runs none. Architecture: gateway, workspace files, heartbeats, skills One long-lived Node gateway per host binds to loopback on port 18789, owns every channel connection, routes inbound messages to sessions, loads context, calls the configured model, executes tools, streams the reply and persists everything under ~/.openclaw. Nodes are paired devices (laptop, phone, headless box) that lend the gateway local screen, camera and shell. Memory is markdown in the agent workspace: AGENTS.md (operating instructions), SOUL.md (persona and boundaries), IDENTITY.md, USER.md (stable facts about you, with its own character budget), MEMORY.md (curated durable facts), a memory/ directory of daily notes, and a one-time BOOTSTRAP.md interview. The memory docs state there is no hidden state; a hybrid memory_search index covers the memory file and daily notes, a background "dreaming" sweep promotes recurring material into MEMORY.md, and a flush runs before context compaction. Initiative comes from the heartbeat, a periodic main-session turn (30 minutes by default, 60 on subscription auth) that can stay silent via a no-reply marker, and from the automations scheduler (one-shot, interval or cron, delivered to a channel, a webhook or nowhere). HEARTBEAT.md is legacy; its checklist now lives in DB-backed scratch. Skills follow Anthropic's Agent Skills format (SKILL.md with name and description front matter) and install from ClawHub; the docs say to treat third-party skills as untrusted code and read them first. The browser tool drives a dedicated agent-owned Chrome/Brave/Edge profile through a loopback-only control service with a strict SSRF policy. The model behind it: hosted versus local OpenClaw is a harness over sixty-plus providers using provider/model references, including Ollama, llama.cpp, LM Studio, vLLM and SGLang for local inference. The tradeoff is the one from the agents episode with higher stakes: everything the agent reads is forwarded to the model, so inbox triage on a hosted model sends your inbox to the provider. Frontier models are better at the judgment calls (is this urgent, is this instruction really from me), local is the only defensible choice for regulated third-party data, and a per-task split (local for reading private content, hosted for writing public content) is a common compromise. The docs make no claim about local-model quality inside OpenClaw. Subscription auth reuses an existing Claude CLI login or an OpenAI OAuth flow per the OAuth docs, which describe the Claude CLI path as sanctioned per Anthropic staff guidance and warn that a community proxy needs a terms check; no published Anthropic term was found either way. Integrations: coding agents from your phone, email, calendar The old Claude Code bridge skill is superseded by agent runtimes: built-in, Codex app-server, Claude CLI and a Copilot plugin, with external harnesses (Claude Code, Gemini CLI, OpenCode, Cursor) driven over the Agent Client Protocol through acpx. Tasks land in managed worktrees: isolated branches with checkpoints in a state DB, filesystem snapshots where supported, a cap around 100 live worktrees, and dirty or unpushed work never auto-cleaned. See Agentic Software Engineering for why worktrees are the right isolation unit. The official IMAP plugin watches a mailbox, spawns an isolated restricted-reader session per allowed message, ranks trust by DMARC/SPF/DKIM, and does not send mail, which encodes the read-versus-write distinction the security section relies on. Calendar and most SaaS arrive as skills or MCP servers; Cisco's DefenseClaw announcement describes connecting email, calendar and Discord through Zapier-hosted MCP servers so a glue service holds the OAuth tokens. Security: what happened in 2026 and what to do about it Exposure: Bitsight, SecurityScorecard and Censys counted between 30,000 and 135,000 internet-facing instances in early 2026, roughly two thirds with no authentication; Censys confirmed 63,070 live instances at the end of March. Bugs: CVE-2026-25253 (CVSS 8.8), a one-click RCE where the control UI auto-connected to a gatewayUrl from the query string and leaked the auth token, worked even against loopback-bound instances and was patched in v2026.1.29 per the GitHub advisory; CVE-2026-32922 (CVSS 9.9) let a pairing token rotate itself into admin, fixed in v2026.3.11. Well over a hundred advisories were logged between February and April. Supply chain: Koi Security's ClawHavoc report found 341 malicious ClawHub skills (335 from one campaign) disguised as wallets, trading bots and Workspace integrations, delivering the Atomic macOS Stealer and targeting always-on Mac minis; the count later passed 800 as the registry grew past 10,000 (The Hacker News, Unit 42). Cisco's skill research scanned about 31,000 agent skills, found a quarter with at least one vulnerability, and demonstrated exfiltration through an attacker-controlled Telegram bot. Injection: Giskard exploited a live deployment for exfiltration and account takeover; PromptArmor showed Telegram and Discord link previews exfiltrate data with no click; CrowdStrike called a misconfigured instance "a powerful AI backdoor agent." Ambient: Wiz found Moltbook's database open with about 1.5 million API tokens; Meta banned OpenClaw on work devices and then acquired Moltbook; China restricted state use. Responses: fast patches, loopback default, DM pairing codes for unknown senders, openclaw security audit, the VirusTotal partnership scanning every ClawHub skill with daily rescans, openclaw skills verify, publisher gating, and the layered access model in the security docs: DM modes, per-agent profiles, control-plane tool restrictions, node exec policies, sandbox and read-only variants, exec approvals, strict browser SSRF, and "one trust boundary per gateway." None of it fixes indirect prompt injection; the maintainers say scanning is not a silver bullet, skills remain arbitrary code, and the agent still holds real credentials. Cisco's open-source DefenseClaw adds pre-execution scanning and runtime allow/block enforcement. Posture, as eight rules: loopback plus a tunnel (Tailscale or SSH) and an auth token even locally; the agent gets its own OS user, mailbox, calendar, browser profile and capped API keys; every integration starts read-only with human approval on irreversible writes; a dedicated box with nothing else on it; read every skill before enabling; design so a hijacked agent is a nuisance, not a breach, by shrinking the write surface; read the logs and memory files weekly; run the audit after every change and pin a version. Use cases that survived By mid-2026 the consumer frenzy had cooled and what remained was solo founders and small teams running it as infrastructure. Surviving uses share one shape, a scheduled or triggered read delivered to the chat you already use: morning briefings, read-only inbox triage with thresholds, watching deadlines and pipelines and even school lunch menus, voice memos returned as structured notes, lead research and CRM updates, and coding-agent dispatch from a phone. Value compounds through the memory files rather than any single automation, which is why the posture has to precede the setup. Alternatives Claude Cowork and OpenAI's ChatGPT Work give the delegate-a-task shape sandboxed and session-based, without messaging or always-on. The rewrites give a smaller, auditable surface. The SDKs from AI Agents i
What an AI agent actually is, why coding agents got good first, how memory really works, what MCP and A2A standardize, which SDKs are alive, how to evaluate on trajectories, and where the products stand after browser agents contracted. Links Try a walking desk - stay healthy & sharp while you learn & code More OCDevel shows - this one has siblings, each on its own subject and produced the same way First of two episodes on AI agents. This one is the architecture: the loop, tools and verifiable feedback, memory, the protocols (MCP, A2A, computer use), the SDK landscape, evaluation and observability, the product map, and when multiple agents help. The next episode, OpenClaw and the Personal Agent, applies it to one always-on assistant with security as the centerpiece. Coding-agent products and mechanics live in the vibe coding sequence starting at MLA 22. Agent vs workflow vs chat: the loop A chat model returns a message; a workflow is your code calling a model at fixed steps; an agent is a model that owns the control flow, choosing its next action from what it observes. That puts systems on a spectrum (chat, chat plus tools, workflows, agents) rather than in a binary, the framing Anthropic's Building Effective Agents uses. The loop itself is ReAct (Yao et al.): thought, action, observation, repeat, with the reasoning trace letting the model track and update a plan. What changed by 2026 is not the loop but the infrastructure around it, and every part of that infrastructure is an attack on per-step error compounding. Tools, function calling, and verifiable feedback Function calling: you describe tools as schemas, the model emits a structured call, your code executes it and returns the observation. The model never runs anything itself, which is the security model. Writing effective tools for agents gives the practical rules: few high-impact tools, clear namespaces, meaningful identifiers, token-efficient responses, descriptions treated as prompt engineering. Effective context engineering for AI agents adds the overlap test: if a human cannot say which tool applies, neither can the agent. The central principle: coding agents got good first because tests and compilers give verifiable feedback that catches a bad step inside the same loop that made it. Find or manufacture the verifier before writing the prompt. Memory: context, retrieval, files, episodic "Memory" means four things: the context window (the only memory the model has), retrieval from an external store, files on disk, and episodic records of prior sessions. Most agent memory is files. Anthropic's memory tool is a client-side file protocol (view, create, replace, insert, delete) against storage you own. The hard part is context management, and both labs converged on the same three mechanisms: context editing to clear stale tool results, compaction to summarize near the limit, and notes written to files before summarization. OpenAI's Responses API conversation state has the same shape with a compaction threshold and compact endpoint. Third-party layers Mem0, Letta (from MemGPT), and Zep (temporal knowledge graph) now compete with first-party primitives. Multi-session patterns: Effective harnesses for long-running agents. Protocols: MCP, A2A, computer use Model Context Protocol is the agent-to-tool standard, now a Linux Foundation project with individual-maintainer governance. 2026 additions: elicitation (server asks the user mid-operation), an extensions mechanism, and the async Tasks extension for long-running tools. Every major SDK below consumes it; its cost is the context each connected server's tool list occupies. A2A is the agent-to-agent standard, Google-built, Linux Foundation-hosted, at v1.0 with a steering committee spanning AWS, Cisco, Google, IBM, Microsoft, Salesforce, SAP and ServiceNow. Strong governance, weak observed consumption; worth knowing, not yet worth building on for small teams. Computer use is the universal fallback: Anthropic's computer use tool (GA toolset with zoom and an automatic injection classifier), Google's Gemini computer use, open-source Browser Use, and Playwright MCP, which drives the accessibility tree instead of screenshots. Prefer API, then accessibility tree, then screenshots. Building one: the SDKs Both labs advise starting without a framework: Building Effective Agents and OpenAI's A Practical Guide to Building Agents. The 2026 SDKs have converged on that critique as thin harnesses around a loop. Claude Agent SDK: Claude Code's loop as a library (built-in tools, subagents, hooks, MCP, permissions, compaction); TypeScript and Python; pre-1.0. OpenAI Agents SDK: handoffs, guardrails, sessions, tracing on the Responses API. OpenAI deprecated the visual Agent Builder in favor of it. LangGraph and LangChain 1.x: stateful graph with checkpointing, interrupts, durable execution; create_agent as a minimal middleware harness. LangSmith is the separate tracing product. Google ADK: code-first hierarchical agent trees with native A2A; deploys to Vertex Agent Engine. Microsoft Agent Framework: GA successor to AutoGen and Semantic Kernel; Python and .NET. CrewAI: role-based crews, past 1.0, with a commercial management platform. smolagents: code agents that write Python instead of JSON calls; weakest maintenance signal on the list. Pydantic AI and Vercel AI SDK: typed validation-first agents in Python; loop control and agent abstraction in TypeScript. Decision rule: machine-operating agent fast, Claude Agent SDK; lightweight handoffs, OpenAI; durable human-in-the-loop state, LangGraph; inside Google or Microsoft, their kit; to understand what you run, write the loop yourself first. Evaluation and observability Agents are evaluated on trajectories, not answers: traces, task evals, cost per task. Traces follow the OpenTelemetry GenAI semantic conventions; products include LangSmith, Langfuse (open source, acquired by ClickHouse), Arize Phoenix, Braintrust, W&B Weave, and Helicone. Public benchmarks show the shape of a task eval: SWE-bench Verified, which OpenAI stopped reporting citing contamination; SWE-bench Pro; tau2-bench; Terminal-Bench 2.0; OSWorld-Verified; GDPval. Cost and reliability: Princeton's Holistic Agent Leaderboard (paper) and its reliability dashboard separate pass@k capability from pass^k reliability; METR time horizons with their own limitations note. Guardrails: OpenAI agent safety, NeMo Guardrails, Guardrails AI; prompt injection framed by Simon Willison's lethal trifecta and Google's CaMeL architectural defense. Products Claude Cowork: "Claude Code for everyone," a sandboxed desktop agent with open-sourced plugins. ChatGPT agent remains; the Atlas browser was retired within a year, folded into ChatGPT and Codex. Google discontinued Project Mariner and moved the capability into Gemini and Antigravity, which absorbed Gemini CLI. Standalone browser agents contracted; the capability moved into models and existing apps. Still shipping: Perplexity Comet (free), Manus (ownership contested this year; check before building on it), Devin, Copilot Studio, Agentforce. Glue: n8n (AI Agent node inside a drawn workflow, MCP server trigger) and Zapier Agents with Zapier MCP. Browser-agent prompt injection is the documented security problem: the PleaseFix research note. Multi-agent: when it helps Two essays a day apart: How we built our multi-agent research system (orchestrator plus parallel subagents beat a single agent on research at roughly 15x the tokens; token usage explained most of the variance) and Cognition's Don't Build Multi-Agents (dispersed decisions and unshared context make it fragile). The disagreement is task shape. Parallelize independent, read-mostly work; keep stateful, sequential work single-threaded; prefer a small hierarchy where workers return findings rather than decisions. Related episodes MLA 29: OpenClaw and the Personal Agent MLA 22: Vibe Coding in 2026 MLA 23: Inside a Coding Agent MLA 24: Agentic Software Engineering Companion show: Agentic Business on Gnothi follows one business as agents take on research, software, sales and operations.
How to automate AI media end to end: clone your own voice on open TTS, pick music that's actually licensed, run ComfyUI graphs headless, design around fal, Replicate and provider queues, finish with ffmpeg, and stay inside licensing at every layer. Links More OCDevel shows - this one has siblings, each on its own subject and produced the same way Companion show. This episode is the overview of the media pipeline. For weekly, hands-on coverage of the video half, from a first usable clip to scenes that cut together, listen to AI Video Generation on Gnothi. Try a walking desk - stay healthy & sharp while you learn & code Pipeline, not prompt Once you need thirty clips with the same character, a narrator who sounds identical every episode, a matched music bed, word-accurate captions and a platform-safe export, the prompt is one node in a graph and the graph is the product. The engineering lives in the edges: how one model's output becomes the next model's input, how failures retry, what each job cost, and whether the run is reproducible next week. Model choice at the nodes is covered in the two sibling episodes; this one covers everything else: voice with ElevenLabs and Qwen3-TTS, licensed music, ComfyUI on your own card, the fal and Replicate APIs, ffmpeg assembly, and licensing. Voice: cloning, open TTS, and consent ElevenLabs remains the reference point (current flagship Eleven v3, character-based pricing). Professional voice cloning is locked to the requester's own voice behind a live voice check, and the terms require consent attestation for any uploaded voice. On the open side, Breeze TTS 2 topped the open-weights column of the Artificial Analysis speech arena in August 2026, ahead of Fish Audio's S2 Pro; the code is Apache 2.0 but the weights are research/non-commercial, so it is not a commercial self-hosting option. The working set for programmers: Qwen3-TTS (Apache 2.0, 0.6B/1.7B, cloning from seconds of reference audio; the preset-speaker variant does not clone, see my Qwen3-TTS voice cloning guide), Chatterbox (MIT, emotion control, watermarked output), and Kokoro (82M parameters, Apache 2.0, faster than real time on CPU). Fish Audio plays both sides with open weights and a cheap hosted API. Quantized Qwen3-TTS runs podcast-length synthesis on CPU-only instances; see Quantized Qwen3-TTS on CPU and the broader open-source TTS roundup. Hosted alternatives for prototyping: OpenAI text-to-speech and Gemini speech generation. Consent is the legal boundary. Tennessee's ELVIS Act added voice to right of publicity; the federal NO FAKES Act cleared Senate Judiciary in June 2026; EU AI Act Article 50 transparency duties apply from 2 August 2026; Denmark is amending copyright law to cover a person's face and voice. Music and sound effects Warner settled with both Suno and Udio; Universal settled with Udio, which became a no-download walled garden; UMG and Sony are still litigating against Suno, whose terms now grant commercial rights rather than ownership to paid subscribers. Eleven Music is trained on licensed data via Merlin and Kobalt deals and cleared for commercial use on self-serve plans, excluding film, TV and larger games; ElevenLabs sound effects are cleared on any paid plan and support loops. Google exposes Lyria and Lyria RealTime through the Gemini API. Open models: ACE-Step 1.5, YuE, Stable Audio Open (community license, best open option for short effects), HeartMuLa, and Meta's MusicGen, which is non-commercial. ComfyUI and local generation ComfyUI is a workflow runtime with a GUI for designing graphs. Comfy raised $30M at a $500M valuation in April 2026 and ships a desktop app, Comfy Cloud, and API nodes that call paid providers from inside a local graph. Programmatic use is the same /prompt endpoint and websocket the front end uses: export the workflow in API format, patch fields, post, poll. Wrappers like comfyui-api and comfy-pack turn a graph into a scalable service. Alternatives: SwarmUI, InvokeAI, the Krita AI plugin. Hardware: full-precision Flux.2 and Qwen-Image do not fit consumer cards; fp8 and GGUF quantization bring them to 16 to 24 GB. For video, the open Wan releases lag the API versions; the 5B variant does 720p on 24 GB (about 8 GB with Comfy offloading), the 14B variant officially wants 80 GB at full precision and needs GGUF to be consumer-viable, and Wan2GP targets low-VRAM cards. Rule: local for iteration, cloud for volume. APIs and aggregators Start with an aggregator, move to a provider API only for a feature or price it lacks. fal is queue-first: submit, get a request ID, poll or webhook, per-output pricing on popular models, per GPU-second for custom deployments. Replicate has the broader catalog beyond image and video, bills per second of compute for open models, packages custom models with Cog, and joined Cloudflare with the same API. RunPod serverless is the raw GPU option for a custom ComfyUI graph. Provider APIs have converged on the same shape: Veo via the Gemini API (billed per output second, audio included), Kling API (post, store task_id, poll /v1/tasks), Runway API, ElevenLabs API. Design rules: every generation is a job in a durable queue keyed on a hash of inputs, model and seed; store the provider's request ID next to your job ID; honor 429 retry-after with a token bucket per provider; persist prompt, seed, inputs and outputs in object storage; route to a second provider on 5xx. Cost per usable second is list price times your rejection rate. Assembly and finishing ffmpeg is the programmer's editor: concat, overlay, sidechain ducking, caption burn-in, crop to 9:16, loudnorm and export. Human-in-the-loop editors: DaVinci Resolve (free version is a real editor; Studio unlocks most Neural Engine features), Descript with its transcript-as-timeline and Underlord assistant, and CapCut for short-form auto-captions, with a caution about its June 2025 terms change. Upscaling: Topaz retired Video AI for the subscription Topaz Video with the Astra model; open-side, SeedVR2 is single-step, runs on 8 GB and plugs into ComfyUI; Real-ESRGAN for clean stills; RIFE for frame interpolation. Captions: WhisperX gives word timestamps within about 50 ms via forced alignment plus diarization, emitting SRT/VTT; generate styled word-pop overlays from its JSON. Delivery: 1080x1920 9:16, H.264/AAC, roughly 10 to 12 Mbps, 30 fps; YouTube recommended upload settings; target -14 LUFS integrated with a -1 dBTP ceiling as the last pipeline step. Licensing across the stack Five layers, and the output is only as clean as the dirtiest node. Weights: Apache/MIT models (Qwen3-TTS, Kokoro, Chatterbox, open Wan) are clean; FLUX.2 dev is non-commercial without a separate license; Stability's community license allows commercial use under a revenue threshold; MusicGen is non-commercial. Output: the Copyright Office holds that prompts alone are not authorship, and the Supreme Court denied cert in Thaler v. Perlmutter in March 2026, so keep evidence of the human selection and editing. Training data: licensed models are the safe path while label suits continue. People: right of publicity, get written consent. Disclosure: YouTube auto-labels via SynthID and C2PA content credentials since May 2026, and labels on Veo and C2PA-stamped content are permanent. Attach credentials and disclose. Two pipelines Social clip (30 s, 9:16, run 200 times): character sheet from an image-editing model stored with prompt and seed -> templated script -> Qwen3-TTS narration with word timings -> per-shot image-to-video jobs via fal keyed on input hash, webhook completion -> cached licensed music bed -> ffmpeg concat, duck, styled captions from timing JSON, crop, loudnorm, 1080x1920 export -> content credentials and disclosure -> human review queue. Narrated explainer (8 min, 16:9, weekly): human-written script (where copyright rests) -> chunked TTS stitched with short silences -> LLM shot list with timestamps tagged diagram/image/video -> deterministic diagrams, styled images, a few video clips upscaled with SeedVR2 -> licensed music and SFX generated once -> Resolve or ffmpeg assembly, captions from narration timings, loudnorm, 1080p/4K export -> title, chapters from the shot list, disclosure, credentials. Both are the same graph with different shot counts and aspect ratios. Related episodes AI Image Generation and Editing in 2026 AI Video Generation in 2026 AI Agents covers agent orchestration of pipelines like these The companion show for the video half of this pipeline is AI Video Generation on Gnothi.
Sora is shut down, Google runs two video models, Kling 3 does lip-synced dialogue, and open-weight MiniMax H3 is what you can actually fine-tune. What a usable clip costs, which models do native audio, how reference consistency works, and why the unit of work is the shot. Links Notes and resources at ocdevel.com/mlg/mla-26 More OCDevel shows - this one has siblings, each on its own subject and produced the same way. Companion show: for weekly, hands-on coverage of the AI video pipeline, from a first usable clip to scenes that cut together, listen to AI Video Generation. Try a walking desk - stay healthy & sharp while you learn & code Second of three episodes on AI media generation, covering Veo, Kling, Runway and MiniMax H3. Four questions: what a usable clip costs, which models generate sound and dialogue natively, how character and shot consistency work now, and where open-weight video fits for a programmer. Ends with the shot-to-scene mental model. What changed since 2025 Native audio is now the baseline at the frontier: Veo 3.1, Kling 3.0, MiniMax H3 and LTX-2.5 sample audio and frames from one model, so lip movement and sound effects land on the right frame. Clips grew from four or five seconds to eight to fifteen, with a few models advertising thirty. Every serious product ships a reference-conditioning feature (Google "ingredients", Kling "elements", Runway references) that holds a character or object across clips. Image-to-video, not text-to-video, is the professional path: lock the first frame with an image model (see AI Image Generation and Editing), then ask the video model to move it. The old "storyteller vs animator" split resolved in favor of the animators. Google: Veo 3.1 and Gemini Omni Google now runs two video models in two places. Veo 3.1 is the developer baseline on the Gemini API and Vertex, in Quality, Fast and Lite tiers, all with native audio; clips are 4, 6 or 8 seconds, and 1080p/4K are upscales of the 8-second clip. It accepts up to three reference images, first-and-last-frame interpolation, and extend in 7-second steps up to 20 times; the Ingredients to Video update added identity consistency, native vertical and 4K upscaling. At Google I/O 2026 Google announced Gemini Omni; Gemini Omni Flash replaced Veo inside the Gemini app and Flow, taking text, image, video or audio as input and supporting conversational video-to-video editing. Per Flow's model matrix it currently tops out at 10 seconds and 720p. All output carries SynthID; the detector portal is still waitlisted. Sora: shut down OpenAI launched Sora 2 on September 30, 2025 with native audio and a free iOS app; it then hit a copyright reckoning over opt-out character use, SAG-AFTRA and Bryan Cranston pushback on likeness, and a court order barring the word "Cameo". In March 2026 OpenAI announced a two-stage shutdown: app and web closed April 26, 2026, API closes September 24, 2026. NBC's reporting attributes it to reallocating compute to coding, reasoning and enterprise; Sora continues only as internal world-model research. It stays in the episode as the case study of a strong model without a business. Kling 3.0 Kuaishou's Kling 3.0 launched globally in March 2026 as a unified image, video and audio model: up to 15 seconds per shot, native 4K, and per the Kling Omni audio guide lip-synced dialogue in five languages with sound effects and ambience generated in the same pass. The control surface is the point: an elements library built from images or short reference video with per-element voice binding, multi-shot generation with continuity, motion transfer from a reference video, motion brush, six-axis camera control, extend and retake. Sold as a credit-based consumer app with commercial rights on paid tiers, plus a first-party API fronted in the West by fal and Replicate. fal's three per-second prices (audio off, audio on, voice control) make the cost of joint audio-video sampling visible. Runway Gen-4.5 and Aleph Runway Gen-4.5 shipped December 1, 2025, briefly topped the Artificial Analysis leaderboard, and was candid about causal reasoning, object permanence and "success bias" failures. Runway's differentiator is editing: Aleph is video-to-video (new angles, relighting, add/remove objects, restyle), and per the Runway API changelog Aleph 2.0 takes 2-30 second inputs with up to five keyframes; Act-Two transfers a filmed performance onto a character. Gen-4.5's release notes do not claim native dialogue or effects. The Runway API now also resells ByteDance Seedance 2.5 (30-second clips, large reference budgets, audio) and Wan 3, and studio deals with Adobe, AMC Networks and Lionsgate anchor the enterprise story. The leaderboard vs the products As of this recording the Artificial Analysis text-to-video arena (blind pairwise human preference) has none of Veo 3.1, Kling 3.0 or Gen-4.5 in its top five: Wan 3.0, Gemini Omni Flash, fal's post-trained MiniMax H3 Max, MiniMax H3, then Seedance 2.0. The headline products win on control, distribution and enterprise fit, not the taste test. Open weights and the second tier Naming matters for Wan: Wan 2.2 is Apache-2.0 open weight (14B MoE needing an 80GB GPU, or a 5B model for a 24GB card) and its GGUF quantizations still trend on Hugging Face; Wan 2.5, 2.7 and 3.0 have no published weights on the Wan-AI Hugging Face org or GitHub and are served as APIs. The open-weight center of gravity is MiniMax H3: 33B, native stereo audio, up to 2K, 4-15 seconds, a community license permitting commercial use, official ComfyUI workflows, and a LoRA and step-distillation ecosystem. LTX-2.5 (19B, native audio, community license free under $10M revenue) and HunyuanVideo-1.5 (8.3B, 14GB with offload, no audio) round out the runnable set; MAGI-2 is a preview with no confirmed license. Elsewhere: Luma shipped Ray3, Ray3 Modify and Ray3.14; Pika pivoted to effects, an agent and MCP on Pika 2.5; ByteDance's Seedance 2.0 and 2.5 ride Dreamina and CapCut distribution; Grok Imagine is a priced API video model outside the top ten; Higgsfield is an aggregator and creative suite, as are fal, Replicate and OpenArt on the developer side. Consistency and control Every consistency feature is conditioning under a different name: text, reference images, first frame, last frame, and reference video are slots the denoiser attends to. A first frame is the strongest condition, which is why image-to-video wins. Reference characters (Veo's three images, Kling elements with voice binding, Seedance's dozens of references) fight identity drift, still the main failure mode. Start and end frames bound a camera move and let shots hand off to each other. Explicit camera controls beat prompt text. Video-to-video (Aleph, Luma Modify, Omni Flash, Kling) means fixing a nearly right shot instead of regenerating. Extend compounds drift, so use it to finish a shot, not build a scene. Open models add LoRAs: Musubi Tuner trains adapters for HunyuanVideo and Wan 2.x, and the tooling lags each new frontier open release by months. Audio in video Native audio means one sampling process produces waveform and frames, conditioned on each other; Veo 3.1 prices everything as video with audio, Kling 3.0 exposes it as a paid toggle, H3 and LTX-2.5 do it in open weights, Gen-4.5 and Wan 2.2 do not. Post-hoc remains a valid choice: MMAudio generates synchronized sound from finished video with an explicit alignment module, and ElevenLabs sound effects generate timed effects from text. Native dialogue holds for a line or two; longer talking heads still favor performance-driven tools like Act-Two. Voice and music proper are in The AI Media Pipeline. Cost per usable second As of this recording, from Gemini API pricing: Veo 3.1 Quality about $0.40/s with audio (720p/1080p), Fast about $0.10/s, Lite about $0.05/s; Gemini Omni Flash is billed per token, working out to roughly $0.10/s of 720p. From Runway API pricing: Gen-4.5 $0.12/s, Aleph 2 $0.28/s with a minimum. From fal: Kling v3 about $0.08/s silent and $0.13/s with audio, Wan 2.5 $0.05/s; Grok Imagine video $0.05-0.08/s. An 8-second Veo Quality shot with audio is a bit over $3; Kling or Gen-4.5 about $1. No vendor publishes success rates; budgeting four generations per usable shot puts a frontier clip with audio at $3-13 and a Fast or open-weight clip under $1. Iterate on the cheap tier, render on the expensive one. Shot to scene Every model generates a shot: one continuous take, one camera, one action, 4-15 seconds. A scene is three to eight shots cut together, and continuity is your job: same references in every shot, first and last frames handing off, the same elements or LoRA, one audio bed over the cut. Storyboard as shots, lock first frames with an image model, iterate cheap, render expensive, fix with video-to-video, assemble in an editor. Assembly, voice, music, ComfyUI and driving it from code are the next episode. Related episodes AI Image Generation and Editing in 2026 The AI Media Pipeline: Voice, Music, ComfyUI, APIs, and Finishing Vibe Coding in 2026 for Grok's coding products




