Session

How to Test AI Features Before Your Users Do

Learn how to build practical evals for LLM features, catch regressions, score quality, and diagnose failures before they reach production.

LLM-powered features often succeed in demos and still fail in production. A prompt change, a model upgrade, or a change to the tools an agent can call can quietly reduce answer quality, break grounding, or introduce risky behavior.

This talk shows engineers and tech leads how to build a practical eval loop for AI features using the same mindset they already bring to testing software. Using an anonymized internal agent that turns technical incident updates into executive summaries as the running example, we will start with a small app that appears to work, then expose three realistic failure modes: unsupported claims, missed key information, and brittle behavior after a change. From there, we will build a lightweight eval suite that turns those failures into repeatable checks.

The demo will cover a compact golden dataset, simple scoring patterns, and a regression workflow for prompt, model, and tool changes. Attendees will leave able to explain the difference between benchmarks and app evals, compare pass/fail, rubric, and pairwise scoring, implement a small but useful eval dataset, and troubleshoot whether failures come from tool calls, generation, or overall system design.

You cannot make every response deterministic. You can make quality visible before your users find the regression.


Audience: Software engineers and technical leads building LLM or RAG features.
Format: 45–60-minute conference talk with live demonstrations. This reusable profile is not the 120-minute workshop version.
Demo: An anonymized knowledge assistant, three failure modes, a golden dataset, scoring, and a regression workflow.
Evidence: Accepted for The Commit Your Code Conference 2026.
Materials: Talk-specific repository, slides, and recording are not yet published.

Ron Dagdag

Microsoft MVP / Research Engineering Manager @ Thomson Reuters

Fort Worth, Texas, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top