Session
Closing the Loop: Building Dual-Pipeline Offline & Online LLM Evaluation Systems
Moving LLMs to production is just the beginning. The real challenge is continuous evaluation. At monday.com, we built a hybrid evaluation architecture that bridges the gap between live production and offline testing. In this session, you’ll learn how our Offline Evals leverage curated, non-production datasets and judges in CI/CD to block regressions before code reaches users, alongside model benchmarking. Then, we’ll explore our Online Evals framework, which monitors live traffic across Quality, Performance, and Cost. Finally, we’ll demonstrate how these online metrics serve as a critical tool for leadership to measure product customer experience, track AI unit economics, and make data-driven architectural decisions.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top