TestTalks.AI - The Testing AI Podcast

When Determinism Dies: Testing and Securing Agentic AI in Production


Listen Later

The pilot episode where two enterprise veterans explain why your old QA pipeline breaks the moment an LLM enters the building.

In this first episode of Ctrl-Alt-Deploy, host Alex Belotsky (CEO of Hesabon/TestSavant.AI) and co-host Kuba Fiatkevich unpack what changes when non-deterministic AI lands in high-stakes software. They argue that the classic assumption of "input A always yields output B" no longer holds, so pipelines built for binary pass/fail states need a new evaluation stage that judges the quality of probabilistic outputs. The conversation moves from LLM chatbots to the 2026 rush toward autonomous agents that wield tools and execute their own code — where a bad chatbot causes reputational damage but a bad agent can delete a database. They dig into treating agents as untrusted "employees," tightening permission scoping at the architectural level, stripping tool permissions the moment a task ends, and red-teaming for prompt injection, exfiltration, and denial-of-wallet attacks. Alex and Kuba also introduce economic and functional health metrics — cost per successful transaction, human intervention rate, benign rejection rate — and describe an adaptive loop where test results feed guardrails that stay inside a compliance and governance framework. The takeaway: reliability is bounded, probable behavior, not perfection, and it belongs to a cross-departmental effort with QA still the final gate.

**In this episode:**

- Why Determinism No Longer Equals Reliability
- Adding An Evaluation Stage To The Testing Pipeline
- Securing Agentic Systems And Permission Scoping
- Agents As Untrusted Employees
- Red Teaming, Prompt Injection And Denial-Of-Wallet Attacks
- Economic And Functional Health Metrics
- Adaptive Guardrails, Compliance And Governance
- The 80/20 Rule And The AI Skills Gap

**Host:** Alex Belotsky (CEO, Hesabon / TestSavant.AI) with co-host Kuba Fiatkevich — enterprise CTO and high-stakes deployment veteran

> "So a chatbot might give you bad advice and might cause some reputational damage, whereas an agent might delete a database."

🎧 *Subscribe to Ctrl-Alt-Deploy on Apple Podcasts, Spotify and YouTube for conversations on building and shipping reliable AI to production, QA, testing and AI governance.*

## Hashtags

#CtrlAltDeploy #AI #SoftwareTesting #QA #QualityEngineering #AgenticAI #LLM #RedTeaming #AISecurity #Guardrails #AIGovernance #TestSavant #DevOps #MLOps

...more
View all episodesView all episodes
Download on the App Store

TestTalks.AI - The Testing AI PodcastBy Automation Cyborg