Agentic AI is moving from demos into real work - but autonomy only matters if teams can trust what the agent is doing. In this episode of AI Cloud Essentials, host Ritu Jyoti sits down with Uma, Nico, and Brandon for a panel-style conversation about what builders need to make agents reliable across specialized domains, long-running workflows, and production environments.
The group unpacks why the most exciting agent use cases are also the hardest to build: agents that take more steps, operate over longer trajectories, and work in specialized settings like drug development, customer support, robotics, and enterprise workflows. That shift raises a new set of infrastructure and tooling questions around online evaluation, offline benchmarking, observability, monitoring, and optimization.
The conversation also explores the emerging agent stack: Weave for prompt, context, and harness engineering; model training and policy improvement through SFT or GRPO; CoreWeave under the hood to run models effectively; and MCP/skills as a way to bring coding agents into the development process itself. This conversation is for AI builders, platform teams, ML engineers, and enterprise leaders trying to move from agent experiments to production-grade systems.
What you'll take away:
Why longer agent trajectories make evaluation and reliability harder
How online evaluation and offline benchmarking work together for agent systems
Why observability is becoming a must-have layer for models, agents, and production AI
How the agent stack is expanding across prompt engineering, context engineering, harness engineering, and model optimization
Why coding agents, MCP, and skills are becoming part of the workflow for everyone - not just developers
Learn how to build agentic AI around evaluation, optimization, and observability - before autonomy becomes a liability.