Session

AI Testing Beyond the Basics: Ensuring Truthful and Reliable Chatbots and Agents

Traditional software testing methods are ineffective when applied to AI-driven chatbots, RAG systems, and agents. Unlike deterministic systems, LLMs generate probabilistic responses, which makes it difficult to verify correctness, consistency and relevance. So, how can we test for truthfulness, bias, hallucinations and real-world usability? In this session, we will move beyond theory to demonstrate practical testing strategies.

Using Testkube and a combination of popular AI testing frameworks, we will show you how to design automated tests and build an agentic AI testbench to continuously assess the performance of your AI workloads across different scenarios. Whether you’re developing a customer support bot, a RAG-powered assistant or an autonomous agent this presentation will provide you with the necessary tools to ensure that your AI workloads are reliable and perform well in practice.

Mario-Leander Reimer

Managing Director, CTO, #CloudNativeNerd @ QAware GmbH

Rosenheim, Germany

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top