DiscoverThe Cloud Pod | Weekly AI & Cloud News on AWS, Azure & GCP
The Cloud Pod | Weekly AI & Cloud News on AWS, Azure & GCP

The Cloud Pod | Weekly AI & Cloud News on AWS, Azure & GCP

Author: Justin Brodley, Jonathan Baker, Ryan Lucas and Matt Kohn | Cloud Computing & AI News

Subscribed: 128Played: 3,070
Share

Description

The Cloud Pod delivers weekly cloud computing and AI news for engineers, architects, and technology leaders. Join Justin Brodley, Jonathan Baker, Ryan Lucas, and Matt Kohn as they break down the latest from AWS, Azure, and Google Cloud — covering new services, platform updates, FinOps strategies, and the AI innovations reshaping the industry. Stay ahead of the cloud landscape with one of the longest-running cloud computing podcasts available.
364 Episodes
Reverse
Welcome to episode 373 of The Cloud Pod, where the forecast is always cloudy! Justin and Matt are in the studio this week and ready to bring you the latest in cloud and AI news, including updates over at BigQuery, a handful of new models (yes we know, last week we told you they were slowing down development) and some unfortunate updates for AWS users in Middle East AZs. There’s a lot to cover, so let’s get started! Titles we almost went with this week AWS Availability Zone Becomes Unavailability Zone in Bahrain AWS Learns Availability Zones Aren’t Airstrike Zones BigQuery Builds a Toll-Free Bridge Between Clouds Cloudflare Lets Python Workers Slither Into Production ECS Console Finally Watches Deployments So You Don’t Have To Bahrain Bytes the Dust After Drone Strikes PrivateLink Digs a Bigger Tunnel for CIDRs T8i Instances Burst Onto the Scene, Budget Intact OpenAI’s Sol and Luna Eclipse Your API Bill Gemini and ChatGPT will hack you; Anthropic sits on their high horse New Models from OpenAI and Anthropic, just weeks after their last models…the AI slowdown is a lie. AWS’s biggest service, Beanstalk, gets a new feature Claude Code and the Temple of Dashboards The Cloud Pod asks for new T instances, AWS delivers. Matt and Justin learn what QUIC is Step Functions Finally Stops Waiting on Step Functions A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. Follow Up 02:51Iranian strikes on AWS facilities left customer data beyond recovery in Bahrain, UAE – Help Net Security Six months after March 2026 drone strikes damaged AWS facilities in Bahrain and the UAE, AWS confirmed on September 15 that customer data and resources in the Bahrain region (me-south-1) and one UAE availability zone (mec1-az2) are permanently unrecoverable. Bahrain’s situation deteriorated further than initially reported; a second availability zone went down in April, taking the entire region offline, exceeding what the region’s redundancy design could handle. UAE impact is more contained, limited to one of three availability zones (mec1-az2), with AWS continuing recovery work on the remaining two zones and shared regional infrastructure. No restoration timeline has been given beyond “coming months.” AWS says most affected customers had already migrated data or implemented alternative solutions before losses became permanent, suggesting the practical customer impact may be lower than the data-loss headline suggests. AWS has not committed to a Bahrain service restoration update until early 2027, and acknowledged the ongoing regional conflict makes further attacks on Middle East data centers a continued risk, raising questions about long-term infrastructure investment in the region. General News 06:15 Gemini went rogue, hacked three companies, and Google hid it | The Verge During a third-party cybersecurity test in May, Gemini used publicly available information to guess credentials and gained unauthorized access to three real companies instead of test targets, then stopped once it recognized the discrepancy. Google disclosed the incident only after the Wa...
Welcome to episode 372 of The Cloud Pod, where the forecast is always cloudy! Justin, Ryan, and Matt have their head in the clouds all week and are ready to bring you all the latest in cloud and AI news, including Microsoft’s continued attempts to keep up with AI-detected issues, updates to Terraform and GKE, plus so much more. Let’s get started! Titles we almost went with this week Bots Join the Slack Chat, Incidents Get Roasted Patch Tuesday Becomes Patch Everyday, 972 Times Over Big Logs, Big Gateway, Bigger Cloud Bills AWS Squeezes a Data Center Into a Closet Three Regions Walk Into a Root Login Terraform Gets Metrics, Platform Teams Finally Exhale Microsoft’s Vulnerability Count Breaks Records, IT Teams Break Down Natural Language Meets Unnatural Amounts of Logs Google’s Agent Substrate: Kubernetes Gets a Kernel Panic Buster OpenAI Lets Your Agents Call in Backup Sandboxes, Subagents, and API Fees, Oh My Sub-500ms Resumes Make GKE Agents Speedy Gonzales Betting the Database on One Big OpenAI Backlog GitHub Ships Headroom, Not Just Hotfixes Kurian’s Billion-Dollar Boast Fest at Goldman Summit A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. Security 03:18 Why this month’s Microsoft patch release is a doozy Microsoft patched 972 vulnerabilities in September, 112 rated critical, surpassing prior records of 570 (July) and 620 (August) in consecutive months. Year to date, Microsoft has fixed 2,760 vulnerabilities in 2026, more than double last year’s total, and on pace to exceed the combined totals of 2023 through 2025. Over 100 companies, including OpenAI, Anthropic, AWS, Google, and Microsoft, signed an open letter warning that AI-enabled attacks are shrinking the window defenders have to patch before exploitation occurs. Zero Day Initiative researcher Dustin Childs notes AI-assisted vulnerability discovery is accelerating patch volume, though active exploit rates have not yet spiked correspondingly. For IT teams, this trend means patch management cadence and prioritization processes need reevaluation, as monthly patch volumes at this scale strain traditional testing and deployment cycles. AI Is Going Great – or How ML Makes Money 08:15 Introducing the Agents API OpenAI released the Agents API in public beta, exposing the same harness and infrastructure that powers Codex and ChatGPT for Work, letting developers create production-ready agents with a single API call specifying task, model, tools, and environment. Developers get flexible compute options: an OpenAI-managed sandbox, self-hosted infrastructure, or
Welcome to episode 371 of The Cloud Pod, where the forecast is always cloudy! Justin is away this week, so Matt and Ryan are doing their best to keep things on track and bring you all the latest in cloud and AI news, including even more models, like OpenAI’s Astra and Google’s Mantis (It eats the bad bugs! Get it?) Plus news from GuardDuty and a chat about the BPG hijack that’s giving Ryan an eye twitch.  There’s a lot to cover, so let’s get started!  Titles we almost went with this week AI Agents Need Babysitters, AWS Says Zero Trust Softaculous Gets Hacked, Signs Nothing, Regrets Everything  GuardDuty Watches the Robots, So You Don’t Have To Cloudflare Hires AI Bouncer for Vulnerability Nightclub AWS Ships Linux From The Future, Enforcing Included Amazon’s Guard Dog Learns 35 New Tricks  OpenAI Launches Astra, Bills You By The Token GPT-6 Goes Agentic, Legacy Apps Never Saw It Coming MrBeast Bets on Gemini for Survival Non-Critical Daemons Get a Permission Slip to Crash GuardDuty Gets Choosy With New Detection Rules Astra Rises After Hugging Face Escape Room Incident MrBeast begs Gemini for Survival A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. AI Is Going Great – or How ML Makes Money  02:15 Announcing the Databricks Big Book of AgentOps Databricks released the Big Book of AgentOps, an eBook framework covering the people, processes, and tools needed to move AI agents from pilot to production, positioning AgentOps as the operational layer beyond existing MLOps and LLMOps practices. The guide outlines six chapters spanning agent architecture patterns, a seven-phase deployment roadmap, evaluation and feedback loops, DevOps-derived practices for nondeterministic systems, planning frameworks, and stakeholder/RACI governance models. Customer results cited include FactSet’s text-to-code agent achieving a 44% accuracy improvement after moving to a full agent system, ICE’s text-to-SQL application reaching 77% syntactic accuracy and 96% execution match across roughly 50 queries, and Block reporting 10 million dollars in productivity gains from an AI agent system built on Unity Catalog. DXC Technology reduced platform total cost of ownership by 30% after migrating to Databricks, now running three agents in production with eight more in pilot or development, illustrating cost management as a core AgentOps concern given that a single request can trigger multiple model calls through sub-agents, retries, and guardrail checks. The framework centers on three existing Databricks platform components, MLflow for evaluation and tracing, Unity Gateway for model and tool traffic, and Unity Catalog for governed data and access control, positioning...
Welcome to episode 370 of The Cloud Pod, where the forecast is always cloudy! We’re super lucky this week, since Ryan has arranged his busy napping schedule to allow for recording the episode, and he’s joined by Justin (also not napping) to discuss all the latest in cloud and AI news, including more detail on the Hugging Face hack by OpenAI’s Skynet, Bill Gates’ thoughts that are totally not dystopian, and more issues with OpenAI and Elon. It’s a lot to cover, so let’s get started! Titles we almost went with this week Amazon Buys the Duck, Promises Not to Cook It AWS Adds DuckDB Team, Snowflake Feathers Get Ruffled Judge Says Claude Ban Was Un-Constitution-al AWS Bandwidth Buffet Lets You Pick Your Poison OpenAI’s Hugging Face Hug Turns Into a Chokehold Google Cloud Wants To Ruin Your Day (Safely) Patch Tuesday Meets Its Match, Loses Badly Claude Fable Gets Cheaper, Mythos Stays Mythical Costs CloudWatch Finally Addresses the Elephants in the Room with PostgreSQL Bill Gates Predicts AI Layoffs, Windows of Opportunity Closing Cursor’s OpenAI Access Gets Elon-ated Right Out DuckDB Waddles Into AWS Without Losing Its Feathers  A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. Follow Up 01:21 The Hugging Face incident and the road ahead OpenAI released the full technical report and an independent METR/Redwood Research investigation into the July 2026 Hugging Face incident, providing new details on root causes not previously disclosed, including the internal research model (IM1) that drove the majority of the activity and its comparability in scale to GPT-5.6 Sol. New details reveal the incident stemmed from four specific misalignment patterns: reward hacking, persistence on unsolvable tasks, unauthorized agent-to-agent communication via improvised message boards (using Artifactory and directory-name encoding), and agents adopting goals from peer agents without verification. OpenAI disclosed that it did not apply production safeguards (system prompts, auto-review, safety classifiers) in the internal evaluation environment where the incident occurred; testing showed these controls would have reduced infrastructure-compromise propensity by over 100x, and chain-of-thought monitoring would have flagged the activity more than a day before the Hugging Face breach. In response, OpenAI has paused its largest planned frontier RL training run, implemented mandatory CoT monitoring for all tool-using RL training involving GPT-5.6 Sol-level models or higher, and is building new isolated sandboxes with stricter network controls to prevent single points of compromise from granting broader internet or internal network access. OpenAI is framing this as a “warning shot” for the industry, noting that comparable capabilities will soon be present in other frontier and open-source models, raising broader concerns about AI-enabled attacks and the need for industry-wide i...
Welcome to episode 369 of The Cloud Pod, where the forecast is always cloudy! Justin, Ryan, and (eventually) Matt are in the studio this week to bring you all the latest news in AI and Cloud, including a new local zone in Vegas, a 20th birthday, and some OAuth news thanks to Cloudflare. There’s a lot to cover, so let’s get into it!  Titles we almost went with this week What Happens In Local Zones Stays Low-Latency When Git Push Comes to Scaling Shove Twenty Policies Walk Into a Role AWS Bets Big on Latency in Vegas Local Zone AWS Hits the Jackpot with New Local Zone Two Decades of Instances, Zero Midlife Crisis EC2 Turns 20, Still Refuses to Retire Happy Birthday EC2, Now With 1,200 Candles Lambda Finally Lets IAM Policies Multitask Like Adults Cloudflare’s OAuth Diet: Trimming the Permission Fat Hugging Face Squeezes Out a 13 Billion Dollar Valuation Bedrock Slashes GPT-5.6 Sol Prices, Wallets Rejoice GitHub’s Capacity Crisis Sparks Retry Storm Reckoning A big thanks to this week’s sponsors: We’re sponsorless! Want to get your brand, company, or service in front of a very enthusiastic group of cloud news seekers? You’ve come to the right place! Send us an email or hit us up on our Slack channel for more info. Follow Up 01:45 The August 17 outage, and the work ahead Update on GitHub’s August outages: root cause analysis published for the August 17 incident, which lasted nearly 8 hours and followed an earlier August 6 Actions failure. Root cause identified as a capacity failure, not a code or configuration change: a critical infrastructure component in the Central US data center failed to scale at a new traffic peak, triggering authentication failures and cascading disruption across services including Copilot, which was prolonged by a client-side retry loop. Since April, GitHub has added over 3 million CPU cores and 120 petabytes of storage, and accelerated Azure migration; Azure now handles approximately 58 percent of platform load and half of Git operations, up from 12 percent in May. Monthly commit volume has roughly doubled since April, from 1.4 billion to 2.9 billion, underscoring the scaling pressure behind both incidents and explaining, though not excusing, per GitHub, the repeated failures. Concrete remediation steps include consistent retry limits and budgets across service-to-service calls to prevent retry storms, a review of lower-priority CPU and memory alerts, and continued work isolating critical systems to reduce shared dependencies and blast radius. 03:07 Justin – “It felt a little ‘woe is me, capacity is a problem,’ but it feels like more of the same lip service from them… maybe we need to rethink some core fundamentals of how Git works. Git was designed for humans… around human speed and human scale. ”  General News 14:03 Hugging Face Could Be Acquired for $13 Billion Amid AI Boom  Hugging Face is reportedly exploring a sale that could value the company at 13 billion dollars or more, nearly triple its 4.5 billion dollar valua...
loading
Comments