Session
AI Testing Beyond the Basics: Ensuring Truthful and Reliable Chatbots and Agents
Traditional software testing methods are ineffective when applied to AI-driven chatbots, RAG systems, and agents. Unlike deterministic systems, LLMs generate probabilistic responses, which makes it difficult to verify correctness, consistency and relevance. So, how can we test for truthfulness, bias, hallucinations and real-world usability? In this session, we will move beyond theory to demonstrate practical testing strategies.
Using Testkube and a combination of popular AI testing frameworks, we will show you how to design automated tests and build an agentic AI testbench to continuously assess the performance of your AI workloads across different scenarios. Whether you’re developing a customer support bot, a RAG-powered assistant or an autonomous agent this presentation will provide you with the necessary tools to ensure that your AI workloads are reliable and perform well in practice.
Mario-Leander Reimer
Managing Director, CTO, #CloudNativeNerd @ QAware GmbH
Rosenheim, Germany
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top