Reliability Enablers

Reliability Enablers

By Ash Patel & Sebastian VietzBusinessTechnology
Download on the App Store

Reliability Enablers episodes

  • #30 Clearing Delusions in Observability (with David Caudill)

    Observability is going through interesting times. David Caudill believes that delusions are getting in the way of our success in this area. He's a senior engineering manager at Capital One, a US-based bank.



    This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit read.srepath.com
    38 min
  • #29 - Reacting to Google's SRE book 2016 (Chapter 1 Part 2)

    Sebastian and I continue our breakdown of notable passages from Chapter 1 of Google's Site Reliability Engineering (2016) book by Betsy Beyer, Jennifer Pettof, Niall Murphy, et al.


    We covered passages like:



    1. Monitoring is one of the primary means by which service owners keep track of a system's health and availability.

    2. Efficient use of resources is important anytime a service cares about money.

    3. Humans add latency, even if a given system experiences more actual failures. A system that can avoid emergencies that require human intervention will have higher availability than a system that requires hands on intervention.

    4. SRE has found that roughly, 70 percent of outages are due to changes in a live system. Best practices in this domain use automation to accomplish implementing progressive rollouts.

    5. Demand forecasting and capacity planning can be viewed as ensuring that there is sufficient capacity and redundancy to serve projected future demand, the required availability.





    This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit read.srepath.com
    32 min
  • #28 - Reacting to Google's SRE Book 2016 (Chapter 1 Part 1)

    Sebastian and I got together to react to and discuss 5 passages from Chapter 1 of Google's Site Reliability Engineering book (2016) by Betsy Beyer, Jennifer Pettof, Niall Murphy, et al.


    We covered passages like:



    1. The sysadmin approach and the accompanying development ops split have a number of disadvantages and pitfalls

    2. Google has chosen to run our systems with a different approach. Our Site Reliability Engineering teams focus on hiring software engineers to run our products

    3. The term DevOps emerged in industry. One could equivalently view SRE as a specific implementation of DevOps with some idiosyncratic extensions.

    4. Google caps operational work for SREs at 50 percent of their time. Their remaining time should be spent using their coding skills on project work.

    5. Product development and SRE teams can enjoy a productive working relationship by eliminating the structural conflict in their respective goals.



    This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit read.srepath.com
    26 min
  • #26 - Growing as a Site Reliability Engineer (Part 2)

    In part 1, we covered the first truth - that you don't grow in your career merely through tenure. That was a simple one.  Let's explore 2 more truths that are somewhat trickier...




    Background music credit: Luna by KaizanBlue



    This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit read.srepath.com
    20 min
  • #25 - DORA and the Pursuit of Engineering Excellence (with Tim Wheeler)

    DORA metrics are a hot topic among technology executives in all kinds of enterprise. But there's more to engineering culture than solely relying on the numbers it goes you. We have a rare treat for you because Ash got Tim Wheeler on the pod. He doesn't do much of social media or podcast episodes. Tim is Director of Engineering Excellence at SquaredUp where he follows the DORA metrics but emphasizes starting conversations around them rather than setting directives.



    This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit read.srepath.com
    38 min
  • #24 - Growing as a Site Reliability Engineer (Part 1)

    How can you grow as an SRE? You've probably thought about your career progression at some point. Ash put together his initial thoughts on this topic. Listen on to learn how he unpacks the first idea of "You don't get promotions with tenure".



    This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit read.srepath.com
    9 min
  • #23 - The Danger of Unreliable Platforms (with Jade Rubick)

    Jade Rubick needs no introduction in the reliability and observability space. He was VP of Engineering at New Relic from 2010 to 2019.


    It was my pleasure to take on his non-obvious ideas on managing expectations with teams, especially platform-based teams. We had a few spicy ideas to dive into.


    We also touched on topics like enhancing engineering practices, DORA metrics, and so much more. Be sure to listen all the way through to learn Jade's amazing insights.



    This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit read.srepath.com
    30 min
  • #22 - How Google does SRE Consulting (with Yury Niño Roa)

    I did not know that Google itself does consulting around its SRE practices. This is not a sponsored episode LOL! I wanted to talk with my SRE friend, Yury Niño Roa, about her drawings and SRE ideas, but we dove into a whole lot more than that. We spoke about her work at Google's PSO office, the antipatterns she's seen, and a whole lot more. Listen in for an engaging conversation.




    You can follow Yury and her amazing drawings via: https://www.linkedin.com/in/yurynino/



    This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit read.srepath.com
    36 min
  • #21 - Better SRE in 2024 is all we can hope for

    Sebastian is back for this episode to help set out direction for 2024. We reflected during the holidays on the problems SREs faced in 2023 in terms of job insecurity, burnout, and "that really shouldn't be my sole job". Sebastian and I talked about what we hope to bring to the community in 2024 to make SREs and SRE teams stronger, happier, and healthier at their work.



    This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit read.srepath.com
    33 min

About Reliability Enablers

From the publisher's feed

Software reliability is a tough topic for engineers in many organizations. The Reliability Enablers (Ash Patel and Sebastian Vietz) know this from experience. Join us as we demystify reliability…

More shows like Reliability Enablers

Kubernetes Podcast from Google by Abdel Sghiouar, Kaslin Fields

Kubernetes Podcast from Google

179 Listeners