Session
Your AI Agent Worked in Dev. Production Had Other Plans
Everything looked good in development. Then real users, real data, and the real production environment showed up.
With nondeterministic software, even saying an agent “works” is surprisingly vague. Different team members may judge the same behavior differently. In production, several variables also change at once: the context the agent receives, the tools it can use, the data it encounters, and the way customers actually use or misuse the product.
The team building an agent naturally tests the scenarios it expects. But engineers are not very good at pretending to be unfamiliar users or imagining how the product may be used months later. LLMs are much better at adopting unfamiliar perspectives.
In this session, I’ll share how we use LLMs to generate unexpected users, edge cases, and usage patterns, then run them against production-like environments. You’ll leave with a practical method for testing beyond your team’s imagination and turning surprising behavior into repeatable evaluations.
Target audience: AI engineers, software engineers, tech leads, architects, and engineering managers building LLM-based products.
Level: Intermediate.
Preferred duration: 30–45 minutes.
Format: Production experience report with practical examples and a demo.
Prerequisites: Basic familiarity with LLM applications or AI agents.
Source: Lessons from building, testing, and operating AI-agent workflows in production.
Haberman Michael
3× Founder & CTO | Building Reliable Software in the AI Era
Tel Aviv, Israel
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top