
Sign up to save your podcasts
Or


Send us Fan Mail
This week Stephen chats with Jamie Allen (Cheif Technologist AWS & SRE @ EPAM Systems) and Adam Kinniburgh (VP Innovation @ SquaredUp) about the concept of a single pane of glass (SPOG) for SRE.
Is it performance art or something actionable? Can alerting replace the need for dashboards? And are metrics drowning in the wake of distributed tracing?
You can find Jamie at:
LinkedIn: https://www.linkedin.com/in/jlallen/
And the Single Pain of Glass article he wrote here: https://medium.com/site-reliability-engineering-leadership/the-single-pain-of-glass-6e42930e966
You can find EPAM at https://www.epam.com/
And you can find the Google Dapper paper here: https://static.googleusercontent.com/media/research.google.com/en//archive/papers/dapper-2010-1.pdf
You can find Adam at:
LinkedIn: https://www.linkedin.com/in/adamkinniburgh/
X: https://twitter.com/adamkinniburgh
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
This week Stephen brings back Kyle Forster from RunWhen to talk about the purple elephant in the room… “AI”.
What makes it GenAI, LLM, Advanced Statistics, or ML? Kyle shares his experience surrounding building AI powered search engines for SRE troubleshooting commands and how to incorporate a (paid) open source community of experts rather than trust AI by itself. They discuss what search looks like under the hood, why GenAI powered chatbots will or won't take over the SaaS industry, how Digital Assistants can be utilised by SREs to increase productivity (hint: giving them to app developers!), how to make informed decisions when purchasing AI products, and *much* more.
You can find Kyle at:
LinkedIn: https://www.linkedin.com/in/kyforster/recent-activity/all/
And you can find out more about RunWhen at:
Website: https://www.runwhen.com/
Product videos: https://www.youtube.com/@whatdoirunwhen
RunWhen Local: https://github.com/runwhen-contrib/runwhen-local (RunWhen Local is an open source troubleshooting cheat sheet that suggests commands from the RunWhen community for all of the namespaces in your cluster - ready to copy & paste)
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
This week Stephen chats with the internet incident librarian herself, Courtney Nash. They explore what Courtney has learned through meta-analysis of the over ten thousands incidents in the Verica Open Incident Database (VOID). They cover why MTTR needs to go in the garbage, joint cognitive systems, the value of looking at near misses and *much* more.
You can check out the VOID here: https://www.thevoid.community/
The two papers mentioned are:
Ironies of Automation by Lisanne Bainbridge: https://queue.acm.org/detail.cfm?id=3380779
Managing the Hidden Costs of Coordination by Laura Maguire: https://ckrybus.com/static/papers/Bainbridge_1983_Automatica.pdf
You can find Courtney at:
LinkedIn: https://www.linkedin.com/in/nashcourtney/
Twitter: https://twitter.com/courtneynash
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
This week Stephen chats with Martin Thwaites from Honeycomb about how developers can leverage observability to understand what they're building better, solve bugs quicker, and have more time for coding. They also discuss OpenTelemetry (the protocol and semantic conventions), manual versus automatic instrumentation, and how keeping every span of trace data is irresponsible.
You can find Martin at:
LinkedIn: https://www.linkedin.com/in/martin-thwaites-ab445120/
X: https://twitter.com/MartinDotNet
And Honeycomb at https://www.honeycomb.io/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
X: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
Observability is a necessary adaptation to make sense of software systems in the Digital Age, but how can we unlock its power for non-engineer stakeholders (such as executives, product owners, etc)? Perhaps we need a layer of abstraction sitting on top of our detailed observability to get the most out of it.
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
This week Stephen chats with former-Google SRE Matt Brown about being on-call. They cover how to up-lift junior engineers so they can be on-call, what a fair on-call schedule looks like, run-books, and much more.
As you heard, Matt believes flexibility is key to a healthy on-call rotation. Matt is exploring ideas for improvements to existing tooling and products in this space and would love to hear from as many listeners as possible with feedback on what they find useful or frustrating with the existing tools they use to support on-call in their teams. You can reach him at [email protected] or schedule a chat via https://zcal.co/mattb/oncall, please don't be shy!
You can also find Matt at:
Website: https://www.mattb.nz/
LinkedIn: https://www.linkedin.com/in/mattbrown/
Mastodon: https://mastodon.nz/@mattb
Twitter: https://twitter.com/xleem
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
The internet is full of people who want to tell you about SRE, DevOps, and Platform Engineering and how different and similar they are... and will give you the impression that these things compete with each other. But do they? And is it a helpful question to ask in the first place?
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
YouTube: https://www.youtube.com/c/SlightReliability
Instagram: https://www.instagram.com/slight_reliability/
TikTok: https://www.tiktok.com/@the_kiwi_sre
Send us Fan Mail
In this episode Amin Astaneh from Certo Modo discusses his experience undertaking an SRE transformation over several years.
Stephen and Amin cover a lot of ground including making ops work visible, measuring toil, the power of calculating the $ value of work, getting developers on-call, the embedded model for SRE, SLOs, culture change, and a whole lot more.
You can find Amin on his company website https://certomodo.io, LinkedIn: https://www.linkedin.com/in/aminastaneh/ and Twitter: https://twitter.com/aastaneh
The books Amin mentions are...
The Practice of Cloud System Administration: https://www.oreilly.com/library/view/practice-of-cloud/9780133478549/
The Phoenix Project: https://www.oreilly.com/library/view/the-phoenix-project/9781457191350/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
Send us Fan Mail
In this episode Stephen Townshend and Sonja Chevre from Tyk discuss making APIs observable, and some anti-patterns to avoid. They cover GraphQL, OpenTelemetry and semantic conventions, correlation IDs, observability pipelines, and much more.
You can find Sonja on LinkedIn: https://www.linkedin.com/in/sonjachevre/ and Twitter: https://twitter.com/SonjaChevre
You can listen to Sonja's KubeCon talk here: https://youtu.be/IkEUJjRBCbo
You can find Tyk's open source gateway here: https://github.com/TykTechnologies/tyk
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
Send us Fan Mail
In this episode Stephen Townshend and Harinder Seera explore how to monitor and manage the cost of cloud. They discuss FinOps as a cultural practice, anti-patterns for implementing in the cloud, keeping cost down through resources, pricing, and architecture... and much more.
You can find Harinder on LinkedIn: https://www.linkedin.com/in/harinderseera/
You can find the official Slight Reliability podcast website at: https://slightreliability.com/
You can find Stephen at:
LinkedIn: https://www.linkedin.com/in/stephentownshend/
Twitter: https://twitter.com/the_kiwi_sre
Instagram: https://www.instagram.com/slight_reliability/
From the publisher's feed

30,707 Listeners

26,165 Listeners

4,831 Listeners

149 Listeners

43 Listeners

179 Listeners

62 Listeners

2 Listeners

142 Listeners

74 Listeners

4 Listeners