DiscoverShip It Weekly - DevOps, SRE, Platform and Cloud Engineering News
Ship It Weekly - DevOps, SRE, Platform and Cloud Engineering News
Claim Ownership

Ship It Weekly - DevOps, SRE, Platform and Cloud Engineering News

Author: Teller's Tech - DevOps, SRE and Cloud Podcast

Subscribed: 10Played: 108
Share

Description

Ship It Weekly is a short, practical recap of what actually matters in DevOps, SRE, cloud infrastructure, and platform engineering.

Each episode, your host Brian Teller walks through the latest outages, releases, tools, and incident writeups, then translates them into “here’s what this means for your systems” instead of just reading headlines. Expect a couple of main stories with context, a quick hit of tools or releases worth bookmarking, and the occasional segment on on-call, burnout, or team culture.

This isn’t a certification prep show or a lab walkthrough. It’s aimed at people who are already working in the space and want to stay sharp without scrolling status pages, cloud updates, and blogs all week. You’ll hear about things like cloud provider incidents, Kubernetes and platform trends, Terraform and infrastructure changes, and real postmortems that are actually worth your time.

Most episodes are 15–30 minutes, so you can catch up on the way to work or between meetings. Every now and then there will be a “special” focused on a big outage or a specific theme, but the default format is simple: what happened, why it matters, and what you might want to do about it in your own environment.

If you’re the person people DM when something is broken in prod, or you’re building the cloud and platform everyone else ships on top of, Ship It Weekly is meant to be in your rotation.

70 Episodes
Reverse
This week on Ship It Weekly: AWS is retiring Amazon DevOps Guru and pointing customers toward CloudWatch and the newer Amazon DevOps Agent. Kubernetes disclosed a vulnerability where StatefulSet and ControllerRevision permissions can allow cross-namespace pod creation under specific conditions. A vulnerability in Undici can let a malicious WebSocket server crash a Node.js process through compressed data. And Cloudflare launched a new CLI as AI agents grow from 25 percent to 48 percent of Wrangler usage.The bigger theme this week is how the systems around our infrastructure are changing. Managed cloud services still have lifecycles that eventually become migration work. Kubernetes authorization can depend on what controllers do with the resources users are allowed to manipulate. Applications acting as clients still process untrusted data. And infrastructure tooling is starting to treat AI agents as first-class users rather than humans who happen to automate commands.In the lightning round: another Kubernetes vulnerability affecting Windows nodes can expose NetNTLMv2 credentials through NTLM coercion. GitHub now supports custom runners for Dependabot version and security updates. And external systems like a CMDB or internal developer portal can push repository properties into GitHub while remaining the source of truth.And the human closer comes from Lorin Hochstein and SRE Weekly. Some availability risks are probably never going away. Resources are finite, networks fail, security controls can affect availability, and production systems have to change. Preventing individual failures still matters, but incident response is part of reliability engineering too. Sometimes improving reliability means getting better at handling the failures you cannot eliminate.LinksAmazon DevOps Guru End of Support - https://tsn.io/GQHN8Kubernetes CVE-2026-2270: Cross-Namespace Pod Creation - https://tsn.io/BsNs8Undici CVE-2026-85024: WebSocket Denial of Service - https://tsn.io/LncLdCloudflare: Introducing the cf CLI - https://tsn.io/Wk3maCloudflare Forge - https://tsn.io/bAPJuLightning RoundKubernetes CVE-2026-76654: Windows NTLM Coercion - https://tsn.io/36Kc3GitHub: Custom Runners for Dependabot - https://tsn.io/tYl9KGitHub: External Custom Properties - https://www.tellerstech.com/go/s-1fd1396d/Human CloserOmnipresent Availability Risks in Cloud Software - https://www.tellerstech.com/go/s-076db9dd/Our LinksThis Week’s On Call Brief - https://tsn.io/fKB9VShip It Weekly - https://tsn.io/NqkdPOn Call Brief - https://tsn.io/Gpz2d
This week on Ship It Weekly: AWS introduced Elastic Beanstalk Cluster Mode, allowing multiple applications to run on shared EKS infrastructure while AWS handles much of the Kubernetes complexity. CrowdSec published how a software supply-chain compromise led to attackers copying roughly 170 private repositories using a stolen OAuth token. A critical Next.js vulnerability in ImageResponse can lead to remote code execution through attacker-controlled SVG data. And Microsoft disrupted EvilTokens, a cybercrime platform linked to more than 12,000 compromised inboxes across 10,000 organizations.The bigger theme this week is what happens after trust has been established. Elastic Beanstalk Cluster Mode puts more infrastructure behind a managed abstraction, but shared infrastructure still means understanding isolation and blast radius. CrowdSec shows how an initial compromise can become a credential problem long after the malicious code is gone. Next.js shows how something as ordinary as generating a social preview image can expose a server-side execution path. And EvilTokens shows how attackers can use valid access to move faster once inside an account.In the lightning round: F5 has a critical BIG-IP APM vulnerability under active exploitation. GitHub Enterprise Cloud can now export an inventory of credentials with enterprise access, including PATs, SSH keys, OAuth tokens, and GitHub App credentials. Zyxel patched a vulnerability affecting GS1900 switches. And Veeam Agent for Microsoft Windows has a privilege-escalation vulnerability that can lead to SYSTEM access.And the human closer comes back to CrowdSec. Removing the malicious package, patching the server, or reimaging the workstation does not necessarily end the incident. If an attacker already stole an OAuth token, cloud credential, SSH key, session, or registry credential, that access can survive long after the original compromise is gone. Containment means understanding not only how the attacker got in, but what they took with themLinksAWS Elastic Beanstalk Cluster Modehttps://tsn.io/1xaV7CrowdSec Supply-Chain Attack Analysishttps://tsn.io/7yq2fNext.js ImageResponse Security Advisoryhttps://tsn.io/8JvHpMicrosoft: Disrupting EvilTokenshttps://tsn.io/DtbC9Microsoft: EvilTokens and Device-Code Phishinghttps://tsn.io/ZzwtDF5 BIG-IP APM CVE-2026-94127https://tsn.io/sFuKWGitHub Enterprise Credential Inventoryhttps://tsn.io/7bpMnZyxel GS1900 Security Advisoryhttps://www.tellerstech.com/go/s-b2595852/Veeam Agent for Microsoft Windows Vulnerabilityhttps://www.tellerstech.com/go/s-166d3119/This Week’s On Call Briefhttps://tsn.io/Nnd8gShip It Weeklyhttps://tsn.io/NqkdPOn Call Briefhttps://tsn.io/Gpz2d
This week on Ship It Weekly: GitHub Actions workflow execution protections are now generally available, giving organizations more control over who and what can trigger individual workflows. Cisco is patching critical vulnerabilities in Secure Email Gateway, including an actively exploited issue that can lead to remote command execution as root. Helm 3 has reached its final minor release and is heading toward end-of-life in February 2027. And GitHub’s ubuntu-latest Actions runner is preparing to move from Ubuntu 24.04 to 26.04.The bigger theme this week is infrastructure that changes even when your code does not. GitHub is making CI execution permissions more explicit, Helm teams now have a defined migration deadline, and the ubuntu-latest transition is a good example of how a completely unchanged workflow can suddenly be running in a different environment. Pinning everything forever is not necessarily the answer. The important part is knowing which dependencies are allowed to move and testing those changes deliberately.In the lightning round: GitHub Actions checks, workflow runs, and statuses will begin following your configured retention period on October 1. GitHub Advanced Security can now enforce configurations from the enterprise level. GitHub added API support for tracking when self-hosted Actions runner versions lose support. And AI Scan for pull requests can now be used without requiring CodeQL default setup.And the human closer starts with a sentence almost every infrastructure engineer has heard during an incident: “But nothing changed.” Maybe nothing changed in the application, but the runner image changed, a dependency moved, a certificate expired, DNS changed, or an external service behaved differently. Latest tags, loose version constraints, external APIs, and even support windows are dependencies. The goal is not to freeze everything forever. It is to avoid accidental mutability, where something can change without the team realizing it was ever allowed to change.LinksGitHub Actions Workflow Execution Protectionshttps://tsn.io/fbqifCisco Secure Email Gateway Security Advisoryhttps://tsn.io/jX2wkHelm 3 End of Lifehttps://tsn.io/Ii7jbUbuntu 26.04 GitHub Actions Runners and ubuntu-latest Migrationhttps://tsn.io/7IJ9kGitHub Actions Retention Changeshttps://tsn.io/idFxyGitHub Advanced Security Configuration Enforcementhttps://tsn.io/8vRMxGitHub Actions Self-Hosted Runner Lifecycle APIhttps://tsn.io/9UhY1GitHub Code Scanning AI Scanhttps://tsn.io/ULAVWThis Week’s On Call Briefhttps://www.tellerstech.com/go/26w38/Ship It Weeklyhttps://tsn.io/NqkdPOn Call Briefhttps://tsn.io/Gpz2d
This week on Ship It Weekly: Amazon Linux 2027 enters public preview with kernel 7.1+, SELinux enforcing by default, DNF5, newer language runtimes, AWS-LC, and an x86-64-v3 baseline. GitHub Actions adds explicit cache permissions to reduce cache-poisoning risk. GitHub can now block pull requests from merging when they introduce exposed secrets. And N-able N-central has a critical pre-auth RCE that Huntress says is being actively exploited in the wild. The bigger theme this week is catching problems before they turn into incidents. Amazon Linux 2027 gives teams time to test AMIs, bootstrap scripts, agents, Terraform, CloudFormation, and CI/CD before the next platform generation becomes production reality. GitHub’s new cache controls make workflow trust boundaries explicit instead of leaving them implied. And secret-scanning rulesets move credential detection directly into the merge path, where developers can actually act on it. In the lightning round: Karmada graduates from the CNCF as multi-cluster and distributed AI scheduling grow, ShieldCrash research claims another Microsoft Defender patch bypass with SYSTEM-level access, CodeQL 2.27 adds native Linux ARM64 support, and Dependabot can now read private GitHub Packages without another personal access token.And the human closer is about what happens when observability shares the same failure domain as the thing it is watching. A full disk is bad enough. It gets worse when logs stop writing, monitoring data disappears, and the tools used to diagnose the outage start failing too. The takeaway is not that every monitoring component needs total isolation. It is that you should know what can blind you, and make sure at least one useful signal survives the failures you care about most.LinksAmazon Linux 2027 Public Previewhttps://tsn.io/NHlEaAmazon Linux 2027 Overview and Preview Detailshttps://tsn.io/izDYxAmazon Linux 2027 Known Issues and Preview Limitationshttps://tsn.io/tdugdGitHub Actions Cache Permissions with cache-modehttps://tsn.io/8p94nBlock Pull Requests with Exposed Secrets from Merginghttps://tsn.io/BspA2N-able N-central 2026.3 Hotfix 4https://tsn.io/xredGHuntress: N-able N-central Vulnerability and Active Exploitationhttps://tsn.io/QjNd5Karmada Graduates from the CNCFhttps://tsn.io/lcJGhMicrosoft Defender ShieldCrash Zero-Day Researchhttps://tsn.io/YfsJ6CodeQL 2.27 Adds Linux ARM64 Supporthttps://tsn.io/lxfnFAutomatic Dependabot Access to GitHub-Hosted Registrieshttps://tsn.io/pSilAShip It Weeklyhttps://www.tellerstech.com/go/siw/On Call Briefhttps://www.tellerstech.com/go/ocb/
This week on Ship It Weekly: AWS Gateway Load Balancer gets TCP Reset, giving applications a faster way to recover when firewalls or other inline appliances fail instead of waiting minutes for TCP retries to time out. Microsoft puts Enterprise Live Migrations into public preview for moving Azure DevOps repositories to GitHub Enterprise Cloud with data residency while developers keep working. GitHub is beginning enforcement against outdated self-hosted Actions runners. And Omarchy fixes a Docker configuration that effectively gave normal desktop processes a path to root.The bigger theme this week is failure modes hiding inside infrastructure we already trust. A dead network path can look like a slow application. A repository migration involves far more than copying Git history. A self-hosted runner can quietly become unsupported while it continues looking healthy. And giving a developer access to the Docker socket may sound like convenience until you remember that the Docker group is effectively a root-level privilege.In the lightning round: Lambda gets full IAM resource-based policies, AWS warns that circular PostgreSQL role memberships can stall major RDS and Aurora upgrades, a researcher releases the FalconFlank CrowdStrike privilege-escalation PoC while CrowdStrike investigates, and SonicWall patches two SMA1000 zero-days after confirming active exploitation.LinksAWS Gateway Load Balancer TCP Resethttps://www.tellerstech.com/go/s-d7e609ab/Azure DevOps Enterprise Live Migrations Public Previewhttps://www.tellerstech.com/go/s-ea05aff9/GitHub Actions Self-Hosted Runner Minimum Version Enforcementhttps://www.tellerstech.com/go/s-6e8540c4/Omarchy: Any User Process Can Escalate to Roothttps://www.tellerstech.com/go/s-d22971c3/AWS Lambda Full IAM Resource-Based Policieshttps://www.tellerstech.com/go/s-ff2a04b5/Fix Circular Role Dependencies Before Upgrading RDS and Aurora PostgreSQLhttps://www.tellerstech.com/go/s-e4578f52/FalconFlank CrowdStrike Privilege Escalation PoChttps://www.tellerstech.com/go/s-8c21b00b/SonicWall SMA1000 Zero-Day Advisoryhttps://www.tellerstech.com/go/s-559ffc8b/Remote Incident Reviews: Async First, Live Later?https://www.tellerstech.com/go/s-68ca9f5e/This Week’s On Call Briefhttps://tsn.io/L95NSShip It Weeklyhttps://www.tellerstech.com/go/siw/On Call Briefhttps://www.tellerstech.com/go/ocb/
loading
Comments