Agent code can look productive right up until a dependency changes, an eval misses the real failure mode, or an over-permissioned tool turns a routine task into a security incident. So what does it actually take to operate AI systems responsibly?
In this episode, Mehdi, Dumky de Wilde, and Maria Vechtomova connect MLOps and LLMOps to agent evals, MCP governance, regenerated software, security, and the engineering practices that still matter when outputs are non-deterministic.
All links and note : https://motherduck.com/podcast
Chapters:
00:00 Meet Maria Vechtomova
01:01 From MLOps to forward-deployed engineering
01:59 Principles first, Databricks second
06:18 What changes from MLOps to LLMOps
09:10 Deterministic tools for non-deterministic systems
09:50 Who maintains regenerated software?
12:38 Hiring for critical thinking with AI
17:38 MCP skills, files, extensions, and stateless servers
20:43 The missing governance layer for agent tools
24:15 Turning deployment pain into reusable practices
28:44 Testing and evaluating LLM systems
31:59 Why AGENTS.md beat skills in Vercel’s evals
40:13 When an agent accidentally hacks Hugging Face
45:00 Skills and the software supply chain
47:53 Consulting that leaves teams stronger
50:42 Wrap-up