
Sign up to save your podcasts
Or


The Gremlin CEO on chaos engineering, resilience and finding failure before your customers do.
Kolton Andrus, who pioneered failure-injection tooling at Amazon and Netflix before founding Gremlin, explains how precise, application-level chaos experiments raised availability while cutting paging by 25%. He argues against separate QA and ops teams as an anti-pattern, insisting the team that writes code should own its quality and resilience. Distinguishing traditional from modern testing, he stresses that distributed systems fail through third-party dependencies, so engineers must test what happens when a dependency fails or slows, building timeouts, fallbacks and graceful degradation. He favors real user monitoring and canary traffic over synthetic data, ties reliability back to brand and revenue, and describes building a shared catalog of failure scenarios so teams can find and fix the low-hanging resilience risks.
**In this episode:**
**Guest:** Kolton Andrus — CEO and Co-founder, Gremlin
> “If you can take a critical failure and turn it into a non-critical failure and someone doesn't have to get woken up in the middle of the night... that's a lot better than just a hard failure.”
🎧 *Subscribe to The QA Lead on Apple Podcasts, Spotify and Google Podcasts for conversations on software testing, QA, test automation and the future of quality engineering.*
## Hashtags
By Automation CyborgThe Gremlin CEO on chaos engineering, resilience and finding failure before your customers do.
Kolton Andrus, who pioneered failure-injection tooling at Amazon and Netflix before founding Gremlin, explains how precise, application-level chaos experiments raised availability while cutting paging by 25%. He argues against separate QA and ops teams as an anti-pattern, insisting the team that writes code should own its quality and resilience. Distinguishing traditional from modern testing, he stresses that distributed systems fail through third-party dependencies, so engineers must test what happens when a dependency fails or slows, building timeouts, fallbacks and graceful degradation. He favors real user monitoring and canary traffic over synthetic data, ties reliability back to brand and revenue, and describes building a shared catalog of failure scenarios so teams can find and fix the low-hanging resilience risks.
**In this episode:**
**Guest:** Kolton Andrus — CEO and Co-founder, Gremlin
> “If you can take a critical failure and turn it into a non-critical failure and someone doesn't have to get woken up in the middle of the night... that's a lot better than just a hard failure.”
🎧 *Subscribe to The QA Lead on Apple Podcasts, Spotify and Google Podcasts for conversations on software testing, QA, test automation and the future of quality engineering.*
## Hashtags