Google Cloud Platform Podcast

The Art of SLOs with Alex Bramley


Listen Later

Today on the podcast, Jon Foust is back with Mark Mirchandani as we talk about SLOs and the importance of measuring service reliability with Alex Bramley. As a member of the Google SRE team, Alex and his coworkers help customers optimally run their services on Google Cloud. They collaborate with the client, weighing client needs and user needs to develop a plan that is affordable, efficient, and has the highest reliability for the user. Recently, they’ve been working to automate functions such as detection of outages, so that Google and the customer can work together quickly to get everything working smoothly again.

Later, Alex, describes the steps developers go through at his workshop, The Art of SLOs, which was designed to help companies measure and improve reliability. At this workshop, attendees are encouraged to set SLO targets and error budgets. They are given theoretical reliability problems to solve, allowing them to practice without the added pressure of messy, real-world problems. The Art of SLOs helps developers understand what measurements are beneficial and why and the best way to implement projects that can take those measurements accurately. Alex was able to make the materials for the workshop free online!

Alex Bramley

Alex Bramley joined Google in January 2010 as the first Mobile SRE in London, after IBM bought the startup he enjoyed working for and made it much less fun. He spent around 7½ years in various reincarnations of Mobile/Android/Play SRE, looking after the infrastructure that makes phones smart, keeps them up to date, and provides them with countless distracting apps.

CRE offered an interesting opportunity to do something different and learn from a bunch of very smart senior people, and Alex has not regretted taking the leap into the unknown. Much of his time recently has been spent rethinking how people teach customers, partners and the general public about SLOs. He helped create the Coursera course on measuring and managing reliability and developed what became the Art of SLOs for Liz Fong-Jones to deliver with other Google SREs at SREcon EMEA’18.

Alex works four days a week so he can (suffer) enjoy looking after his children on Wednesdays, listen to cheerful music, and waste a lot of time playing computer games and occasionally writing code.

Cool things of the week
  • Postponing Google Cloud Next ’20: Digital Connect blog
  • New research: How effective is basic account hygiene at preventing hijacking blog
  • Simplified global game management: Introducing Game Servers blog
Interview
  • The Art of SLOs site
  • CRE Life Lessons blog
  • Putting customers first with SLIs and SLOs blog
  • Putting customers first with SLIs and SLOs (Part 2) blog
  • Measuring and Managing Reliability course
  • Site Reliability Engineering books
Question of the week

How do I get started with GCGS? docs

  • Google for Games Developer Summit Keynote video
  • Google for Games Developer Summit Playlists videos
Where can you find us next?

Jon will be working on an Open Match sample project for the developer community.

Mark will be making more videos like Error Reporting and error logging - Stack Doctor.

...more
View all episodesView all episodes
Download on the App Store

Google Cloud Platform PodcastBy Google Cloud Platform

  • 4.6
  • 4.6
  • 4.6
  • 4.6
  • 4.6

4.6

101 ratings


More shows like Google Cloud Platform Podcast

View all
The Vergecast by The Verge

The Vergecast

3,664 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

624 Listeners

Acquired by Ben Gilbert and David Rosenthal

Acquired

4,224 Listeners

AWS Podcast by Amazon Web Services

AWS Podcast

202 Listeners

The Daily by The New York Times

The Daily

110,759 Listeners

Kubernetes Podcast from Google by Abdel Sghiouar, Kaslin Fields

Kubernetes Podcast from Google

184 Listeners

Talks at Google by Talks at Google

Talks at Google

118 Listeners

The Journal. by The Wall Street Journal & Spotify Studios

The Journal.

5,957 Listeners

Google DeepMind: The Podcast by Hannah Fry

Google DeepMind: The Podcast

196 Listeners

Hard Fork by The New York Times

Hard Fork

5,448 Listeners

Huberman Lab by Scicomm Media

Huberman Lab

28,602 Listeners

Cloud Security Podcast by Google by Anton Chuvakin

Cloud Security Podcast by Google

38 Listeners

The Weekly Show with Jon Stewart by Comedy Central

The Weekly Show with Jon Stewart

10,459 Listeners

Google Cloud Basics by Jason Meers

Google Cloud Basics

0 Listeners

The AI Daily Brief (Formerly The AI Breakdown): Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief (Formerly The AI Breakdown): Artificial Intelligence News and Analysis

504 Listeners