Session

Production Ready RAG with Ragas framework

Building a RAG application takes a weekend; making it production-ready takes months. For engineers, the biggest bottleneck is evaluation: relying on manual spot-checking is slow, unscalable, and makes tracking regressions impossible.This deeply technical session covers how to treat LLM evaluation like software engineering. We will dive into the architecture of Ragas, exploring how it uses LLMs-as-a-Judge to quantify system performance. We will break down the math and logic behind the RAG Triad such as Faithfulness, Answer Relevance, and Context Precision/Recall. You will learn how to bootstrap your testing with automated, evolutionary synthetic test data generation. Finally, we will demonstrate how to embed Ragas directly into your GitHub Actions or CI/CD pipelines to automatically block hallucinations before code hits production.

Audience Takeaways

Algorithmic Evaluation: Understand the internal prompt-logic and scoring mathematics Ragas uses to calculate metrics.

Isolating the Failure Point: Learn to pinpoint whether a bad user response is a retrieval failure (vector DB) or a generation failure (LLM).

Automated Test Generation: Master the "evolutionary query" method to automatically synthesize diverse, production-grade test datasets from your raw documents.

Continuous Integration for AI: Walk away with an architectural blueprint to build an automated, metric-driven evaluation step inside your CI/CD pipeline.

Swathy Santhoshkumar

Forward Deployed Engineer-Senior Manager at Accenture

Tampa, Florida, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top