Anuj Tyagi is a site reliability engineering (SRE) leader specializing in AI infrastructure, with over a decade of experience building and scaling large-scale cloud-native platforms, including the past year focused specifically on agentic AI infrastructure. He is a co-founder of AITechNav, a nonprofit mentoring people in AI, SRE, and cloud engineering, and is recognized as an AWS Community Builder, an IBM Champion, and a Platform Community Ambassador.
Anuj has spoken at conferences including HashiConf, DevOpsDays, and DevConf.US, is an active open-source contributor to CNCF and other projects, writes on his dev.to blog and on LinkedIn, and describes himself on the show as an Alibaba Cloud Community Builder as well.
In This Episode...
Everyone is building something with AI right now, but very few of those projects reach production. In this episode, host Ghazenfer Mansoor, CEO of Technology Rivers and a podcast host known for AI, SaaS, and HIPAA-compliant HealthTech development, talks with Anuj Tyagi, a site reliability engineering leader whose years in AI infrastructure give him a clear view of exactly where that gap comes from. The conversation surfaces five recurring failure points that keep AI projects stuck in proof-of-concept: skipping caching and timeout handling, overloading agents with too many MCP tools, launching without guardrails, ignoring RAG faithfulness metrics, and building without a fallback gateway.
Anuj explains the shift from probabilistic LLM output to the deterministic results real products need, and how Model Context Protocol (MCP) gives agentic AI infrastructure real access to outside tools and data. He also covers how teams keep AI costs under control with caching, guardrails, and gateway fallback, and breaks down RAG hallucination detection: how hallucinations happen, why confident-sounding wrong answers are so dangerous, and which metrics catch them early.
Zooming out, the episode makes a case for engineering discipline as the real differentiator in the AI era. Prompting alone will not produce a reliable system. The teams that make it to production are the ones treating site reliability engineering for AI as a first-class problem, not an afterthought.