Session

Testing Agentic Systems: How to Evaluate AI Workflows That Don’t Behave Like Traditional Software

Traditional software testing assumes that the same input should lead to the same output. Agentic AI systems break that assumption. An LLM-powered workflow may vary its wording, choose a different reasoning path, or behave differently depending on context, tool results, or prompt phrasing. That makes evaluation one of the hardest parts of building real AI applications.
This talk presents a practical, software-engineering-focused approach to testing agentic systems in production. We will look at how to move beyond prompt demos and start validating AI workflows more like real software systems: with scenario-based tests, structured output checks, routing validation, adversarial inputs, and failure-mode evaluation. We will also explore what should still be tested deterministically, what should be evaluated behaviorally, and how to think about acceptable variability in systems that are not fully deterministic.
The focus throughout is on practical architecture and engineering tradeoffs, not model research. Attendees will leave with a clearer framework for evaluating LLM-powered applications, improving reliability over time, and building greater confidence before and after release.

Key takeaways
- A practical framework for testing AI workflows that include nondeterministic model behavior.
- Concrete approaches for validating routing, structured outputs, safety checks, and failure handling.
- A production-minded way to think about reliability, regression testing, and quality in agentic systems.

Target audience
Software developers, architects, tech leads, QA-minded engineers, and AI engineers who are building or planning LLM-powered applications and want practical ways to evaluate them beyond ad hoc prompting.


Suggested tags:
AI engineering, testing, agentic AI, LLMs, software quality, production systems

Compact Version:
Traditional testing assumes deterministic behavior. Agentic AI systems do not behave that way. LLM-powered workflows may vary in wording, reasoning path, or decisions depending on context, tool results, or prompt phrasing, which makes evaluation much harder than in traditional software.
This talk presents a practical approach to testing agentic systems using scenario-based tests, structured output checks, routing validation, adversarial inputs, and failure-mode evaluation. Attendees will leave with a framework for improving reliability and confidence in AI-powered applications before and after release.

Eyal Wirsansky

Staff AI Engineer | Adjunct AI Professor | Author of ‘Hands-On Genetic Algorithms with Python’ | JUG and GDG Community Leader

Jacksonville, Florida, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top