
Sign up to save your podcasts
Or


Google’s SRE Book popularized the concept of Service Level Objective (SLO) and the SLO-driven approach. But what does it really mean to make SLO driven decisions? How can we generate observability and synchronize teams around joint SLOs? And how can we automate SLOs and integrate them into the software release pipeline?
In this episode I’ll host Andreas Grabner. We’ll discuss the SRE practices, and how to automate SLO from dev all the way to prod. We’ll talk about the open source efforts to standardize the process under the Continuous Delivery Foundation, and about Keptn, the new CNCF open source project that promises to help with this automation.
Andreas Grabner (@grabnerandi) has 20+ years of experience as a software developer, tester and architect and is an advocate for high-performing cloud scale applications. He is a contributor and DevRel for the CNCF open source project keptn (www.keptn.sh). Andreas is also a regular contributor to the DevOps community, a frequent speaker at technology conferences and regularly publishes articles on blog.dynatrace.com or medium. In his spare time you can most likely find him on one of the salsa dancefloors of the world.
The episode was live-streamed on 15 March 2022 and the video is available at https://youtu.be/J81byOpVqrk
OpenObservability Talks episodes are released monthly, on the last Thursday of each month and are available for listening on your favorite podcast app and on YouTube.
We live-stream the episodes on Twitch and YouTube Live - tune in to see us live, and pitch in with your comments and questions on the live chat.
https://www.twitch.tv/openobservability
https://www.youtube.com/channel/UCLKOtaBdQAJVRJqhJDuOlPg
Show Notes:
Resources:
Socials:
What does it take to build observability in a web-scale company such as Slack, Pinterest and Twitter?
On this episode of OpenObsevability Talks I'll host Suman Karumuri to hear how he built these systems from the ground up on these #BigTech co's, about his recent research papers and more.
Suman Karumuri is a Sr. Staff Software Engineer and the tech lead for Observability at Slack. Suman Karumuri is an expert in distributed tracing and was a tech lead of Zipkin and a co-author of OpenTracing standard, a Linux Foundation project via the CNCF. Previously, Suman Karumuri has spent several years building and operating petabyte scale log search, distributed tracing and metrics systems at Pinterest, Twitter and Amazon. In his spare time, he enjoys board games, hiking and playing with his kids.
The episode was live-streamed on 16 February 2022 and the video is available at https://youtu.be/IvidkV3TfYg
OpenObservability Talks episodes are released monthly, on the last Thursday of each month and are available for listening on your favorite podcast app and on YouTube.
We live-stream the episodes on Twitch and YouTube Live - tune in to see us live, and pitch in with your comments and questions on the live chat.
https://www.twitch.tv/openobservability
https://www.youtube.com/channel/UCLKOtaBdQAJVRJqhJDuOlPg
Show Notes:
* Who owns observability in large organizations?
* The gaps in current way of handling metrics
* MACH research paper for metrics storage engine
* The gaps in current way of handling logs Slack KalDB
* SlackTrace - Slack in house tracing system
Resources:
Socials:
SaaS (software as a service) is a popular model for many businesses today. SaaS businesses need agility to move fast and remain competitive. This means agility in the software IT stack, but also agility in the business models and product-led growth (PLG). Observability plays a key role in enabling SaaS organizations to move fast.
Achieving this agility, however, raises specific observability requirements. On this episode of OpenObservability Talks we’ll host Aviad Mizrachi, the CTO and Co-Founder of Frontegg, to help us map these requirements. Having escorted dozens of SaaS businesses across many verticals, Aviad brings a wealth of experience in how today’s SaaS is built and operated, and will share his insights and best practices on how to design and build the observability stack right.
Aviad has been a developer for the last 20 years. He held a few management and architecture positions on startups such as Vicon and HTS as well as in larger companies such as NICE and CheckPoint. Today at Frontegg Aviad works closely with many customers to help them build their SaaS solutions.
The episode was live-streamed on YouTube Live and Twitch on 11 Jan 2022 and the video is available at https://www.youtube.com/watch?v=ZcneTMeBPeg
OpenObservability Talks episodes are released monthly, on the last Thursday of each month and are available for listening on your favorite podcast app and on YouTube.
We live-stream the episodes on Twitch and YouTube Live - tune in to see us live, and pitch in with your comments and questions on the live chat.
https://www.twitch.tv/openobservability
https://www.youtube.com/channel/UCLKOtaBdQAJVRJqhJDuOlPg
Show Notes:
Resources:
Socials:
We’ve grown to rely on “the three pillars” for observability - logs, metrics and traces. Popular frameworks such as Prometheus have helped popularize these practices. But now people are starting to realize that it’s not enough.
On this episode Dotan Horovits will host Frederic Branczyk for a discussion about the unspoken pitfalls of Prometheus and the challenges of current observability coverage. We will also discuss the rise of Continuous Profiling as a new observability signal, what it’s about and where it can help. We’ll also review the recent launch of Parca, an open source project for continuous profiling that traces its roots to Red Hat’s internal ConProf open source tool.
Frederic is the founder and CEO of Polar Signals. Before founding Polar Signals he was a senior principal engineer and the main architect for all things Observability at Red Hat, which he joined through the CoreOS acquisition. Frederic is a Prometheus and Thanos maintainer as well as the tech lead for the special interest group for instrumentation in Kubernetes. In a previous life, he was a security researcher working on key management solutions as well as intrusion detection systems. When not working on software Frederic enjoys obsessing over brewing a perfect cup of coffee.
The episode was live-streamed on 16 December 2021 and the video is available at https://www.youtube.com/watch?v=G02g63oI0IA
OpenObservability Talks episodes are released monthly, on the last Thursday of each month. The episodes are also live-streamed on Twitch and YouTube Live - tune in to see us live, and pitch in with your comments and questions on the live chat.
Show Notes:
Resources:
Social:
OpenObservability Talks S2E06: Hosting Steve McCanne
We hear a lot about BPF in the industry today, applying this flexible technology to solve so many problems from routing, proxying, and of course observability. Correlating events and data from the operating system level across distributed systems is a key problem for the industry and community to solve. I am thrilled to announce Steve McCanne joining us for this episode. I have been lucky enough to spend time with Steve in my career and am delighted to have him join us to discuss the origin stories and where these foundational technologies might be applied in the future. Steve’s Bio and background speak for themselves.
Steve McCanne is the "Coding CEO" at Brim, a small startup working on the open-source Zed Project and a new application called "Brim" that leverages Zed. Back in the days before the Web, Steve worked at the Lawrence Berkeley National Laboratory where he developed BPF, libpcap, the PCAP file format, and the tcpdump language and compiler, while also working on the Real-time Transport Protocol (RTP) for Internet video when the telcos claimed that real-time Internet communication was impossible without end-to-end virtual-circuit guarantees. (Guess who was right?) After a brief stint in academia in the late '90s, Steve crossed over to the dark side, became a tech entrepreneur, and never looked back. He has founded several startups and took his '02 company and Sharkfest's sponsor, Riverbed, public in '06.
Resources
The USENEX paper from 1993 on BPF architecture: https://www.usenix.org/legacy/publications/library/proceedings/sd93/mccanne.pdf
Open source tools Steve shared in the podcast:
https://github.com/brimdata/zed
Steve's GitHub BPF repo: https://github.com/brimdata/zbpf
Twitter: https://twitter.com/OpenObserv
YouTube: https://www.youtube.com/@openobservabilitytalks
Have you ever wondered how services are operated at Google’s scale? Here’s your opportunity to find out. Ramón will share how his SRE team runs Google’s identity services, and the elaborate end-to-end observability they use to achieve it with strict SLA. We’ll also get a glimpse at the birthplace of Kubernetes, OpenCensus, Dapper, Monarch and other cornerstones of today’s cloud-native DevOps and observability.
Ramón Medrano Llamas (@rmedranollamas) is a staff site reliability engineer at Google, focused on user identity and authentication. He concentrates on the reliability aspects of new Google products and new features of existing products, ensuring that they meet the same high bar as every other Google service. Before joining Google in 2013, he worked at CERN developing and designing distributed systems for physics. He holds a master’s degree in computer science and is pursuing a PhD on distributed systems.
The episode was live-streamed on 26 October 2021 and the video is available at https://youtube.com/live/jVTZf1SXZrg
Show Notes:
Resources:
Observability is becoming a common practice for DevOps teams monitoring and troubleshooting IT systems. But Observability can offer much more than that. More advanced usage of telemetry, and in particular distributed tracing and its context propagation mechanism, can uncover insights into your business performance and can help solve business and FinOps problems.
On this episode of OpenObservability Talks I hosted Yuri Shkuro, creator of Jaeger project and a champion of Distributed Tracing, to discuss how tracing and observability can help beyond DevOps, whether on business cases, FinOps or even software development. We also caught up on the latest updates from Jaeger, the CNCF’s distributed tracing OSS project, its synergy with OpenTelemetry and more topics.
Yuri is a software engineer who works on distributed tracing, observability, reliability, and performance problems; author of the book "Mastering Distributed Tracing"; creator of Jaeger, an open source distributed tracing platform and a graduated CNCF project; co-founder of the OpenTracing and OpenTelemetry CNCF projects; member of the W3C Distributed Tracing Working Group.
The episode was live-streamed on 14 September 2021 and the video is available at https://youtube.com/live/YSOyTagKGtM
Show Notes:
Resource:
In this episode, we’ll talk with industry veteran and product manager Anurag Gupta who has been working in open source observability for over 4 years. We will go into depth on his background, and how he views the ecosystem of open source. Then we will dig into the Fluentd and Fluent Bit projects and discuss some of the amazing innovations coming from this project. Learn what’s next for logging, and how a consolidated data collection plane is being driven by the Fluentd project.
The CNCF has a rich suite to address monitoring Kubernetes and cloud-native workloads. First of which is Prometheus, which is widely adopted, with great out-of-the-box compatibility with Kubernetes. But under the CNCF you can also find OpenMetrics that offers standardization of the metrics format, Thanos and Cortex which offer long-term storage for Prometheus, and other complimentary solutions and integrations.
On this episode of OpenObservability Talks we’ll host “RichiH” Hartmann and discuss the different OSS projects, the synergy between them, and the future roadmap in building the community and making CNCF a leading offering.
Richard "RichiH" Hartmann is Director of Community at Grafana Labs, Prometheus team member, OpenMetrics founder, CNCF SIG Observability chair, and other things. He also organizes various conferences, including FOSDEM, DENOG, DebConf, and Chaos Communication Congress. In the past, he made mainframe databases work, ISP backbones run, and built a datacenter from scratch.
The episode was live-streamed on 02 July 2021 and the video is available at https://youtube.com/live/j3nFFHSosnI
Show Notes:
Resources:
Current observability practice is largely based on manual instrumentation, which creates a barrier to entry for many wishing to implement observability in their environment. This is especially true in Kubernetes environments and microservices architecture.
eBPF (extended Berkeley Packet Filter) is an exciting new technology for Linux kernel level instrumentation, which bears the promise of no-code instrumentation and easier observability into Kubernetes environments (alongside other benefits for networking and security).
On this episode of OpenObservability Talks we’ll host Natalie Serrino, Principal Engineer at Pixie Labs, which was recently acquired by New Relic. We’ll talk about observability in Kubernetes environments, eBPF and its use cases for observability.
We’ll also talk about Pixie, the Kubernetes-native in-cluster observability platform, and the exciting news of it being open sourced and contributed these days to CNCF under Apache 2.0 license.
Natalie is a Principal Engineer and Tech Lead at New Relic. She works on the Pixie auto-telemetry observability platform, which was acquired and open sourced by New Relic. She focuses primarily on Pixie’s data layer, including its query language, compiler, and query execution engine.
The episode was live-streamed on 20 June 2021 and the video is available at https://youtube.com/live/NYDBj5ctKaw
Show Notes:
Resources:
http://www.brendangregg.com/ebpf.html
https://blog.px.dev/
https://docs.px.dev/about-pixie/roadmap/
https://www.businesswire.com/news/home/20210504005480/en/New-Relic-Joins-Cloud-Native-Computing-Foundation-Governing-Board-and-is-in-the-Process-of-Contributing-Pixie-Open-Source-for-Kubernetes-Native-Observability
https://netflixtechblog.com/how-netflix-uses-ebpf-flow-logs-at-scale-for-network-insight-e3ea997dca96
https://logz.io/blog/istio-instrumenting-microservices-distributed-tracing/
https://opensearch.org/blog/update/2021/06/opensearch-release-candidate-announcement/
https://thenewstack.io/tracing-why-logs-arent-enough-to-debug-your-microservices/
https://www.theregister.com/2021/06/29/kubernetes_spend_report/
Socials:
Twitter: https://twitter.com/OpenObserv
YouTube: https://www.youtube.com/@openobservabilitytalks
From the publisher's feed