
Sign up to save your podcasts
Or


When one input can return a hundred different answers, "does it pass?" stops being a yes-or-no question — and testing has to be reinvented.
SAP AI advisor Ian McCallum — who literally wrote the book on testing SAP applications — joins Alex and Kuba to unpack why generative AI and agents have shattered deterministic testing. The conversation moves from why enterprises deploy AI to how they do it safely: governance, security, and lifecycle management. Ian argues that regression testing is dead for probabilistic systems, replaced by statistical and spot testing borrowed from life sciences — sampling, bell curves, and allowable variances tuned to how tight "good enough" needs to be (very tight for closing the books, looser for some supply chains). They dig into the coming swarm of agents, why "red teaming" is the wrong frame for LLM quality, and how the human-in-the-loop role shifts from clicking submit to evaluating results across a defense-in-depth stack. The episode closes on a genuinely unsettling question: what happens when an intelligent agent learns to hide its own problems?
**In this episode:**
**Guest:** Ian McCallum — AI Technical Solution & Strategy Advisor, SAP
> "It's not a bug. It's a feature. As I keep saying."
🎧 *Subscribe to Ctrl-Alt-Deploy on Apple Podcasts, Spotify and YouTube for conversations on building and shipping reliable AI to production, QA, testing and AI governance.*
## Hashtags
By Automation CyborgWhen one input can return a hundred different answers, "does it pass?" stops being a yes-or-no question — and testing has to be reinvented.
SAP AI advisor Ian McCallum — who literally wrote the book on testing SAP applications — joins Alex and Kuba to unpack why generative AI and agents have shattered deterministic testing. The conversation moves from why enterprises deploy AI to how they do it safely: governance, security, and lifecycle management. Ian argues that regression testing is dead for probabilistic systems, replaced by statistical and spot testing borrowed from life sciences — sampling, bell curves, and allowable variances tuned to how tight "good enough" needs to be (very tight for closing the books, looser for some supply chains). They dig into the coming swarm of agents, why "red teaming" is the wrong frame for LLM quality, and how the human-in-the-loop role shifts from clicking submit to evaluating results across a defense-in-depth stack. The episode closes on a genuinely unsettling question: what happens when an intelligent agent learns to hide its own problems?
**In this episode:**
**Guest:** Ian McCallum — AI Technical Solution & Strategy Advisor, SAP
> "It's not a bug. It's a feature. As I keep saying."
🎧 *Subscribe to Ctrl-Alt-Deploy on Apple Podcasts, Spotify and YouTube for conversations on building and shipping reliable AI to production, QA, testing and AI governance.*
## Hashtags