Your monitoring dashboard says everything is fine. Your system collapses anyway. As companies like Uber and Netflix went from running single servers to orchestrating 50+ microservices across the globe, traditional alerts became useless—like diagnosing a patient by examining their elbow while the problem festers in the pancreas. Discover how software engineers borrowed a concept from Control Theory to finally see what's really happening inside their systems.
00:00 - The healthy patient who stops breathing
02:30 - Why dashboards can't see around corners
05:00 - One server versus a world tour
08:15 - Kalman's Control Theory meets software debugging
11:45 - Distributed tracing: the x-ray machine for code
15:20 - Why observability changed everything
---
Sources & further reading:
• Benjamin Sigelman et al. — "Dapper, a Large-Scale Distributed Systems Tracing Infrastructure" (Google Technical Report, April 2010): https://research.google/pubs/dapper-a-large-scale-distributed-systems-tracing-infrastructure/
• Cindy Sridharan — Distributed Systems Observability (O'Reilly, July 2018): https://www.oreilly.com/library/view/distributed-systems-observability/9781492033431/
• CNCF — "Cloud Native Computing Foundation Announces OpenTelemetry's Graduation" (May 21, 2026): https://www.cncf.io/announcements/2026/05/21/cloud-native-computing-foundation-announces-opentelemetrys-graduation-solidifying-status-as-the-de-facto-observability-standard/
• Peter Bourgon — "Metrics, tracing, and logging" (February 21, 2017): https://peter.bourgon.org/blog/2017/02/21/metrics-tracing-and-logging.html
• Charity Majors / Honeycomb — "Observability: A Manifesto": https://www.honeycomb.io/blog/observability-a-manifesto
• Charity Majors / Honeycomb — "Observability 5-Year Retrospective": https://www.honeycomb.io/blog/observability-5-year-retrospective
• CNCF — "OpenTelemetry has graduated… Now what?" (July 24, 2026): https://www.cncf.io/blog/2026/07/24/opentelemetry-has-graduated-now-what/
• CNCF — "A Brief History of OpenTelemetry (So Far)" (May 2019): https://www.cncf.io/blog/2019/05/21/a-brief-history-of-opentelemetry-so-far/
• Grafana Labs — "OpenTelemetry: Challenges, priorities, adoption patterns, and solutions": https://grafana.com/opentelemetry-report/
• Elastic — "Observability trends for 2026: Maturity, cost control, and driving business value": https://www.elastic.co/blog/2026-observability-trends-costs-business-impact
• MonitoringCost.com — "Observability Cost as % of Cloud Spend: 7–12% Median (2026)": https://monitoringcost.com/observability-cost-as-percent-of-cloud
• Grepr.ai — "What Are The Hidden Costs in Observability in 2026?": https://www.grepr.ai/blog/the-hidden-cost-in-observability
• OneUptime — "The True Cost of Observability Tool Sprawl 2026": https://oneuptime.com/blog/post/2026-02-28-true-cost-of-observability-tool-sprawl/view
• Cloudflare — "Understanding how Facebook disappeared from the Internet" (October 2021): https://blog.cloudflare.com/october-2021-facebook-outage/
• Broadcom — "From Kálmán to Kubernetes: A History of Observability in IT": https://academy.broadcom.com/blog/aiops/from-kalman-to-kubernetes-a-history-of-observability-in-it
• Google Open Source Blog — "OpenTelemetry: The Merger of OpenCensus and OpenTracing" (May 2019): https://opensource.googleblog.com/2019/05/opentelemetry-merger-of-opencensus-and.html
• W3C — "Trace Context" (W3C Recommendation, 2020): https://www.w3.org/TR/trace-context/
• Uber Engineering Blog — "Evolving Distributed Tracing at Uber Engineering": https://www.uber.com/blog/distributed-tracing/
• SoundCloud Backstage Blog — "Prometheus Monitoring at SoundCloud": https://developers.soundcloud.com/blog/prometheus-monitoring-at-soundcloud/
• Rudolf E. Kalman — "On the general theory of control systems" (IFAC Proceedings, 1960)
• …and 1 more in the episode research notes
This podcast episode was fully generated by AI — research, script, voices, and production. Built with Claude, Piper TTS, and automated pipeline tooling.