Session
Don’t Trust, Test: Practical Evaluation in Microsoft Foundry
Generative AI should not reach production based only on impressive demos or answers that “look right.” Just as code needs unit tests, models, prompts, and agents need a structured evaluation process.
In this session, we will explore how to use Microsoft Foundry to evaluate models and agents, starting from the fundamentals and moving toward practical test automation scenarios. We will look at how to define evaluation datasets, choose meaningful metrics, compare results, detect regressions, and integrate evaluation into a more robust development workflow.
The goal is to show how evaluation can evolve from a manual, occasional activity into a continuous engineering practice for building more reliable, measurable, and production-ready AI solutions.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top