Session
Your AI System Passed the Demo. How Do You Know It Works?
Traditional software gives us a familiar question: does the test pass?
AI systems make that question much harder. The same input can produce different outputs, a response can be factually correct but poorly grounded, and an agent can complete a workflow successfully while making an unsafe intermediate decision.
This session explores how to evaluate AI applications beyond a handful of example prompts. We will build an evaluation approach around correctness, relevance, groundedness, tool behaviour and failure cases, and examine how evaluations can become part of the development workflow rather than a final check before production.
Through practical examples, we will compare manual testing, curated datasets and automated evaluation, and discuss where each approach breaks down. We will also look at how evaluation results can be connected to model, prompt, retrieval and agent changes so teams can tell whether an AI system is actually improving.
The goal is simple: replace "the demo looked good" with measurable evidence that an AI system is becoming more reliable.
Monica R
Software Development Engineer @ Autodesk - Speaks AI, Tech & Careers
Bengaluru, India
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top