
Sign up to save your podcasts
Or


Google's Tech Incident Response Team (IRT) manages the complex production environment during times of intensive change. In this episode, IRT member Adam Kramer talks about the importance of psychological safety to ensure that engineers can communicate effectively.
Courtney Nash of The VOID discusses the role of human expertise in managing complex systems, and how SREs continue to bring critical value even as technology and AI evolve.
John Allspaw joins Prodcast hosts Matt Siegler and Florian Rathgeber for a candid discussion of reliability topics at SREcon Americas 2026.
John Allspaw discusses reliability with Prodcast host Steve McGhee at SREcon Americas 2026
We speak with Ricard Bejarano about being an SRE at home, discussing Home Lab systems.
We sit down with Matt Zelesko, VP of SRE at Google, for a candid talk about how AI is changing SRE — and how it's not.
Sam Anderson shares his experiences with burnout, and how to support yourself as a reliable system. Sam provides guidance on how to deal with burnout, and some suggestions on how to avoid burnout through understanding yourself and finding the help and support you need.
Crisis Engineer Mikey Dickerson joins us to talk about what constitutes a crisis. Mikey draws on his broad experience across industry and the public sector, as well as on work with his team of systems fixers.
What's happening in the world of SRE and resilience engineering? Join us as we catch up with fellow podcast hosts Colette Alexander and Clint Byrum of the This Is Fine! podcast at SREcon in Seattle.
How do you introduce Site Reliability Engineering to an AI research lab, bringing concepts of scale to engineers who are at the leading edge of AI systems?
In the latest episode of The Prodcast, hosts Steve McGhee and Florian Rathgeber chat with Damion Yates, who helped establish the reliability engineering culture at Google DeepMind. Damion shares his journey of bringing scalable infrastructure to DeepMind, supporting massive machine learning experiments.
Discover the unique challenges of supporting AI research, such as managing highly expensive "lockstep" training models where a single machine failure halts the entire process. Damion also explains why he believes "luck is our enemy" in systems engineering, and why protecting a research scientist's time is the ultimate metric for success.
From the publisher's feed

32,046 Listeners

30,701 Listeners

43,359 Listeners

286 Listeners

148 Listeners

759 Listeners

684 Listeners

213 Listeners

9,537 Listeners

179 Listeners

1,075 Listeners

565 Listeners

5,557 Listeners

185 Listeners

16 Listeners