Session
Evaluation Lessons for Agents in Production
Agentic applications are everywhere— from autonomous workflows to multi-step AI systems that look production-ready in demos. Once these agents hit real users, real data, and real scale, teams discover an uncomfortable truth: most agentic apps fail silently in production. When something goes wrong, it’s hard to know why, where, or how often.
We will explore a gap between prototype and production for agentic systems: lack of observability and evaluation. This talk breaks down how teams can move beyond “it worked once” to systems that are measurable, debuggable, and reliable.
We’ll walk through architectural patterns for tracing agent decisions, evaluating agent behavior over time, detecting drift and failures, and correlating cost, latency, and security at task level. You’ll see how production-grade observability and evaluation transform agentic apps—from opaque black boxes into systems you can confidently operate, scale, and improve.
This session is not about adding more agents or smarter prompts—it’s about knowing whether your agentic app is actually ready for production, and what to fix when it isn’t.
Anannya Roy Chowdhury
Gen AI Developer/Advocate at Amazon Web Services (AWS)
Bengaluru, India
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top