Session
A Thousand Models in the Catalog, and You Still Picked Wrong
Every AI project starts with the same question: which model? It is the wrong question, or at best the fourth one to ask. In a year of shipping LLM features to production, the model name mattered less than everything around it: whether the step needed a model at all, how the deployment was hosted and versioned, and whether anything measured when answers got worse.
This session is a decision framework with receipts. A document pipeline where the model never sees a number that matters. A chat feature pinned to a model that is now deprecated, with token prices frozen in source code. An eval judge that itself became a deployment problem.
These decisions live in Microsoft Foundry: a catalog of thousands of models, and every way to mis-deploy them. We walk what matters there: deployment types and quotas, data-zone versus global hosting, keys versus Entra-only access, version pinning and upgrade policies, and the eval harness that makes swapping models boring.
The checklist you leave with: decide where a model belongs, pick the smallest one that passes your evals, host it like it will be deprecated, because it will.
For developers and architects taking LLM features to production.
Which model? Wrong question. Where does a model belong at all, how do you host it in Microsoft Foundry, and what tells you when answers get worse? A production-tested framework with receipts: deprecated pins, prices frozen in source, judge models, and evals that make swaps safe.
Nikos Delis
Senior Cloud & Software Engineer | Microsoft MVP for Azure & IoT
Malmö, Sweden
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top