Session

Taming Rogue Agents: Observability-Driven Evaluation for Production Reliability

Moving from a viral demo to reliable enterprise AI is the hardest hurdle in modern software engineering. Because agents are inherently non-deterministic, validating performance requires a granular look at the "why" and "how" behind every action.

This session explores how observability redefines agentic evaluation beyond binary pass/fail testing, shifting the focus from the result to the entire reasoning chain. Featuring a live demo using Arize Phoenix and LangGraph, we’ll showcase how to trace complex loops and automate evaluations in real-time.

What you’ll learn:

Beyond Output: Why judging final answers alone is a recipe for silent failure.

Live Implementation: Using Arize Phoenix with LangGraph to visualize traces and debug agentic "thought processes".

The Playbook: An incremental framework for evaluation-first development.

Are you ready to stop guessing and start measuring? Join us to master the art of agent evaluation.

Apoorva Jaiswal

Vice President - Applied AI ML Lead at JPMorgan Chase & Co.

Palo Alto, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top