KubeFM

KubeFM

By KubeFMTechnology
Download on the App Store

KubeFM episodes

  • Building Platforms for AI Agents, with Mauricio (Salaboy) Salatino

    Non-deterministic agents pose specific challenges for platform teams in observability, state management, governance, and trust.

    Mauricio (Salaboy) Salatino explains why agentic applications behave like distributed multi-agent systems. His test assigned agents to take an order, cook the pizza, deliver it, and charge the customer. One order crossed 15 containers and produced 200 traces.

    In this interview:

    • Why agent frameworks can recreate monolith scaling and resource contention

    • How OpenTelemetry data can measure agent behavior and trust over time

    • Why narrowly scoped agents are safer than agents that follow long sequences

    • How the platform can become the learning layer that feeds better context back to LLMs

    Sponsor

    Kubernetes moves too fast to track everything. Learn Kubernetes Weekly filters out the noise to deliver one curated email with useful articles, tutorials, tools, jobs, events, and CFPs. Subscribe to Learn Kubernetes Weekly.

    More info

    • Find all the links and info for this episode here: https://ku.bz/TlVjXdnb6

    • Interested in sponsoring an episode? Learn more.

    29 min
  • Kubernetes Can Run Your Database. Your Team Can't., with Kat Cosgrove

    KubeSelect is a new show that tests one bold Kubernetes hypothesis with one expert. Hosts Salman Iqbal and Bart Farrell examine the evidence and ask whether the claim holds up.

    In episode two, Kat Cosgrove tackles a long-running question: should teams run databases on Kubernetes? StatefulSets, persistent volumes, CSI, and database operators changed the technical answer, but they did not remove the operational risk.

    In this episode:

    • How Kubernetes storage has matured and why old assumptions persist

    • Where Kubernetes responsibility ends, and database responsibility begins

    • Why operators do not replace database experts

    • When managed database services remain the better choice

    Sponsor

    Kubernetes moves too fast to track everything. Learn Kubernetes Weekly filters out the noise to deliver one curated email with useful articles, tutorials, tools, jobs, events, and CFPs. Subscribe to Learn Kubernetes Weekly.

    More info

    • Find all the links and info for this episode here: https://ku.bz/7yDWlP8T5

    • Interested in sponsoring an episode? Learn more.

    42 min
  • Observability Won't Save You, with Henrik Rexed

    Kube Select takes one bold Kubernetes hypothesis and tests it with an expert.

    In the first episode, Salman Iqbal and Bart Farrell ask whether most organizations have an observability problem or a decision-making problem.

    Henrik Rexed, Senior Staff Engineer at Dynatrace, challenges the original claim. Teams can collect metrics, logs, and traces, but that telemetry needs system relationships, ownership, and deployment history to support a confident decision.

    In this episode:

    • Why telemetry without context slows incident response

    • How SLOs, ownership, and deployment events guide troubleshooting

    • When more metrics increase cost without improving decisions

    • How AI agents can investigate incidents without replacing human judgment

    Sponsor

    Kubernetes moves too fast to track everything. Learn Kubernetes Weekly filters out the noise to deliver one curated email with useful articles, tutorials, tools, jobs, events, and CFPs. Subscribe to Learn Kubernetes Weekly.

    More info

    • Find all the links and info for this episode here: https://ku.bz/848pmBN5_

    • Interested in sponsoring an episode? Learn more.

    32 min
  • Why Kubernetes Needs to Learn GPUs, with Saiyam Pathak

    Kube Signals starts where the keynote ends: with the trends that platform teams will have to operationalize next.

    In this special episode, Brian Teller speaks with Saiyam Pathak about his KubeCon India keynote and the shift from developer platforms to AI factories. They examine what GPU scarcity, shared accelerators, and AI workloads mean after the conference slides meet real infrastructure.

    In this interview:

    • Why GPU infrastructure is becoming a platform-engineering concern

    • How DRA, HAMI, MIG, and MPS change GPU allocation and utilization

    • Where isolation, scheduling, and observability become harder for AI platforms

    • Which cloud-native AI trends and projects platform engineers need to watch

    Sponsor

    This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.

    More info

    • Find all the links and info for this episode here: https://ku.bz/4QZDqrnf-

    • Interested in sponsoring an episode? Learn more.

    33 min
  • GitOps at Enterprise Scale, with Elad Cohen

    At enterprise scale, a deployment pipeline that runs Helm upgrades directly against Kubernetes hides drift, mixes configuration with CI logic, and makes the last pipeline run the source of truth.

    Elad Cohen explains how WSC Sports moved from Azure DevOps to GitHub Actions and redesigned delivery around Git and Argo CD. The resulting platform separates builds from deployments, keeps service configuration in values files, and continuously reconciles clusters.

    In this interview:

    • Why CI should change Git instead of the cluster

    • How ApplicationSets create main and shadow deployments from one values file

    • How AppProjects scope permissions and route alerts by team

    • Why reusable Helm contracts make customization compound across services

    Sponsor

    This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.

    More info

    • Find all the links and info for this episode here: https://ku.bz/wX5H5Mjwv

    • Interested in sponsoring an episode? Learn more.

    37 min
  • Automating Pod Disruption Budgets with Kyverno, with Ahmad Asmar

    Karpenter can reduce Kubernetes infrastructure costs, but aggressive node consolidation can also expose workloads that lack disruption safeguards.

    Ahmad Asmar explains how Zencity uses Kyverno to automatically generate Pod Disruption Budgets, while accounting for existing PDBs, percentage-based availability targets, single-replica workloads, and environment-specific policies.

    In this interview:

    • How Karpenter consolidation changes the availability risks of cluster operations

    • Why Kyverno's generated policies can provide safer defaults than manual enforcement

    • How to handle duplicate PDBs, scaling workloads, and single-replica edge cases

    • How aggregated ClusterRoles keep custom permissions separate from Helm-managed resources

    Sponsor

    This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.

    More info

    • Find all the links and info for this episode here: https://ku.bz/xrlPJg54D

    • Interested in sponsoring an episode? Learn more.

    46 min
  • From KIAM to EKS Pod Identities, with Fabián Sellés Rosa

    An unmaintained identity component can remain invisible until a routine Kubernetes upgrade turns it into an incident.

    Fabián Sellés Rosa, Platform Engineer and Runtime Tech Lead at Adevinta, explains how his team moved from KIAM to EKS Pod Identities without discarding the security boundaries and application interface that their internal platform depended on.

    In this interview:

    • Why KIAM became urgent to replace after years of stable operation

    • How Crossplane, a custom controller, and KRO with ACK compared against the team's criteria

    • Why managed EKS Capabilities reduced toil but introduced observability and rollout trade-offs

    • How Kyverno preserved namespace-level authorization for IAM roles

    Sponsor

    This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.

    More info

    • Find all the links and info for this episode here: https://ku.bz/R_06hwnCn

    • Interested in sponsoring an episode? Learn more.

    27 min
  • 1 Million Tokens Per Second on Kubernetes, with Federico Iezzi

    GPU inference throughput depends on more than accelerator generation or count.

    Memory bandwidth, model parallelism, cache configuration, and the load generator itself all influence measured throughput.

    Federico Iezzi, Customer Engineer at Google Cloud, explains how his team achieved 1 million output tokens per second using Qwen 3.5 27B, vLLM, GKE Autopilot, and NVIDIA B200 GPUs.

    The discussion covers:

    1. Why memory bandwidth limits decode performance

    2. How Federico chose between tensor and data parallelism

    3. What changed after enabling multi-token prediction and reducing the KV cache footprint with FP8 quantization.

    Sponsor

    This episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.

    More info

    • Find all the links and info for this episode here: https://ku.bz/1xD9Md0mb

    • Interested in sponsoring an episode? Learn more.

    48 min
  • The Hidden Cost of Slow Autoscaling, with John Ford

    Forced platform migrations are usually treated as something to survive. At Scout24, a mandatory OS migration became an opportunity to rethink Kubernetes autoscaling, node provisioning, and infrastructure efficiency.

    John Ford explains how Scout24 moved its EKS-based Infinity platform from a polling autoscaler and over-provisioned capacity to Karpenter and Bottlerocket. The result was faster node startup, a safer migration path, and about a 30% infrastructure reduction without major downtime.

    In this interview:

    • Why two-minute node provisioning forced a 25% capacity buffer

    • How Karpenter made the Bottlerocket migration safer

    • What broke around EC2 metadata, AWS SDKs, and cgroups

    • How the new foundation enables Spot, ARM, and GPU workloads

    Sponsor

    This episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.

    More info

    • Find all the links and info for this episode here: https://ku.bz/DdmVC2_7v

    • Interested in sponsoring an episode? Learn more.

    22 min
  • The Namespaces Scaling Trap, with Brian Stack

    Most teams scale Kubernetes by thinking about pods and nodes. At Render, Brian Stack ran into a different dimension: hundreds of thousands of namespaces per cluster, multiplied across DaemonSets that list-watch every namespace.

    Brian explains how Render traced the issue through Calico and Vector, worked with upstream maintainers, and turned memory profiling into operational wins: lower node costs, lighter API-server load, and faster rollouts.

    In this interview:

    • Why namespaces can become a hidden scaling bottleneck

    • How DaemonSets multiply memory and control-plane pressure

    • How profiling, staging clusters, and upstream collaboration freed 7 TiB

    • Why pushing from an 80% fix to a complete fix can make teams faster

    Sponsor

    This episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.

    More info

    • Find all the links and info for this episode here: https://ku.bz/0mrvCsXrV

    • Interested in sponsoring an episode? Learn more.

    37 min

About KubeFM

From the publisher's feed

Discover all the great things happening in the world of Kubernetes, learn (controversial) opinions from the experts and explore the successes (and failures) of running Kubernetes at scale.

More shows like KubeFM

Software Engineering Radio - the podcast for professional software developers by team@se-radio.net (SE-Radio Team)

Software Engineering Radio - the podcast for professional software developers

274 Listeners

The Changelog: Software Development, Open Source by Changelog Media

The Changelog: Software Development, Open Source

286 Listeners

Security Now (Audio) by TWiT

Security Now (Audio)

2,010 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

622 Listeners

LINUX Unplugged by Jupiter Broadcasting

LINUX Unplugged

272 Listeners

The Enterprise AI Show by Massive Studios

The Enterprise AI Show

149 Listeners

Talk Python To Me by Michael Kennedy

Talk Python To Me

581 Listeners

Soft Skills Engineering by Jamison Dance and Dave Smith

Soft Skills Engineering

286 Listeners

Thoughtworks Technology Podcast by Thoughtworks

Thoughtworks Technology Podcast

43 Listeners

Late Night Linux by The Late Night Linux Family

Late Night Linux

169 Listeners

Kubernetes Podcast from Google by Abdel Sghiouar, Kaslin Fields

Kubernetes Podcast from Google

179 Listeners

AWS Podcast by Amazon Web Services

AWS Podcast

201 Listeners

The Stack Overflow Podcast by The Stack Overflow Podcast

The Stack Overflow Podcast

62 Listeners

2.5 Admins by The Late Night Linux Family

2.5 Admins

98 Listeners

Oxide and Friends by Oxide Computer Company

Oxide and Friends

66 Listeners