Google SRE Prodcast

Google SRE Prodcast

By Salim VirjiTechnology
Download on the App Store

Google SRE Prodcast episodes

  • The One with STPA, Jeffrey Snover, and Theo Klein

    This episode discusses Systems Theoretic Process Analysis (STPA), a method for analyzing complex systems. Theo Klein, a Google SRE, and Jeffrey Snover, a Distinguished Engineer at Google, explain that STPA focuses on identifying how system accidents and losses occur due to a loss of control, rather than component failures. STPA helps identify design flaws early, even before code is written! The discussion highlights that STPA is a human-driven process, prompting critical questions about system goals and potential losses, and that Google is adapting the pure STPA approach for commercial software development to make it more practical and efficient.

    38 min
  • The One with Startups and Adam Fletcher

    In this episode, hosts Steve McGhee and Matt Siegler are joined by guest, Adam Fletcher, CEO and Co-Founder of MarketStreet. They discuss the current state of web development with LLMs, managing technical debt in startups, the evolution of infrastructure and reliability engineering, the role of community in technology, and the future of software engineering with AI.

    42 min
  • The One with SLOs and Sal Furino

    In this episode, Sal Furino, Customer Reliability Engineer at Bloomberg, discusses all things Service Level Objectives (SLOs) with hosts Steve McGhee and Matt Siegler. Together, they dig into what successful SLOs look like, how it relates to users, and how SLOs provide an effective framework for joint decisions about system reliability across product, engineering, and leadership teams.

    44 min
  • The One With the Future of SRE and Matt Zelesko

    Matt Zelesko, the head of Site Reliability Engineering at Google, discusses the evolution of SRE, highlighting the shift from traditional operations to a model that balances velocity and reliability to better serve the rapid advancements in AI and ML. He emphasizes that SRE's core mission is to enable partners to move quickly while meeting reliability goals, and that the sheer scale of Google's infrastructure necessitates the SRE model for cross-system problem-solving. Zelesko envisions AI as a crucial assistant for SREs, improving incident detection, mitigation, and postmortem processes, and allowing SREs to focus on more complex engineering challenges and risk management earlier in the development cycle, while still valuing the hands-on experience of operating production infrastructure.

    27 min
  • The One with AI and Todd Underwood

    In this Google Prodcast episode, Todd Underwood, a reliability expert from Anthropic with experience at Google and OpenAI, discusses the current state and future of AI in SRE. Todd and the hosts focus on the current state and future of AI and ML in production, particularly for SREs. Topics discussed include the challenges of AI-Ops, limitations of current anomaly detection, the potential for AI in config authoring and troubleshooting, trade-offs between product velocity and reliability, the evolving role of SREs in an AI-driven world, and book publication for optimal timing.

    44 min
  • The One With Data Centers and Peter Pellerzi

    This episode features guest, Peter Pellerzi (Distinguished Engineer, Google). Peter and the hosts, Matt Siegler and Steve McGhee, focus on the physical infrastructure side of SRE, discussing topics such as the scale of Google's data centers, handling incidents like power outages, testing and preparedness strategies, the use of AI for optimizing cooling plants, and more. Peter also emphasizes the importance of community support, proactive planning, and learning from real-world testing and incidents to ensure high availability and resilience in data center operations.

    37 min
  • The One With Security and Jessica Theodat

    Jessica Theodat (Senior SRE & Security Tech Lead, Google) joins hosts Jordan Greenberg and Steve McGhee to discuss the intersection of security and site reliability engineering at Google. Jessica touches on risk management, the unique nature of security incident responses, and the shared goals between security and SRE. The crew also delves into the balance between security and SRE, acknowledging the tension and the need for collaboration between teams to achieve business goals and user trust.

    20 min
  • We're back with Season 4!

    In this "bumpisode", hosts and producers of Prodcast (including our new co-host, Matt Siegler!) reflect on the previous season and introduce the new season's focus on upcoming trends in Site Reliability Engineering (SRE) and AI, and the friends we make along the way. They also introduce new elements we are bringing in with Season 4, such as a video format and a feedback form.

    16 min
  • Special Episode: You Missed a Page from Telebot

    This episode features Javi Beltran, a Google engineering lead who created the "Telebot" theme song. With our beloved hosts, Steve McGhee and Jordan Greenberg, Beltran discusses the origins of the song, created in 2012 for Google's paging system. The song was meant to add a touch of levity to what could be a stressful situation for engineers on-call. Beltran also unveils a new, more modern remix of "Telebot" (created in collaboration with our host, Jordan Greenberg!) which will be used as the intro theme for the podcast's next season.

    17 min
  • Imperative vs. Declarative Change Workflows with Dominic Hutton & Niccolo' Cascarano

    In this episode of the Prodcast, guests Dominic Hutton (Staff SRE, HashiCorp) and Niccolo' Cascarano (Senior Staff SRE at Google) join hosts Steve McGhee and Jordan Greenberg to dive into configurations. They discuss the differences between imperative and declarative configuration, explore the benefits and challenges of each approach, and the need for careful consideration when choosing between the two. Ultimately, the goal is to achieve reliable and maintainable systems through effective configuration management.

    37 min

About Google SRE Prodcast

From the publisher's feed

SRE Prodcast brings Google's experience with Site Reliability Engineering together with special guests and exciting topics to discuss the present and future of reliable production engineering!

More shows like Google SRE Prodcast

Freakonomics Radio by Freakonomics Radio + Stitcher

Freakonomics Radio

32,046 Listeners

Planet Money by NPR

Planet Money

30,701 Listeners

Hidden Brain by Hidden Brain, Shankar Vedantam

Hidden Brain

43,359 Listeners

The Changelog: Software Development, Open Source by Changelog Media

The Changelog: Software Development, Open Source

286 Listeners

The Enterprise AI Show by Massive Studios

The Enterprise AI Show

148 Listeners

All In The Mind by ABC Australia

All In The Mind

759 Listeners

Warriors Plus Minus: A show about the Golden State Warriors by Audacy

Warriors Plus Minus: A show about the Golden State Warriors

684 Listeners

Python Bytes by Michael Kennedy and Calvin Hendryx-Parker

Python Bytes

213 Listeners

The Indicator from Planet Money by NPR

The Indicator from Planet Money

9,537 Listeners

Kubernetes Podcast from Google by Abdel Sghiouar, Kaslin Fields

Kubernetes Podcast from Google

179 Listeners

The World in Brief from The Economist by The Economist

The World in Brief from The Economist

1,075 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

565 Listeners

Hard Fork by The New York Times

Hard Fork

5,557 Listeners

The Rest Is Money by Goalhanger

The Rest Is Money

185 Listeners

ThursdAI - The top AI news from the past week by From Weights & Biases, Join AI Evangelist Alex Volkov and a panel of experts to cover everything important that happened in the world of AI from the past week

ThursdAI - The top AI news from the past week

16 Listeners