Session
Tracing the Failure: Debugging a Production RAG System in a Language Your Tools Don’t Speak
RAG systems are increasingly used to build question-answering applications over domain-specific data. However, when a RAG application produces a wrong answer, identifying where the failure occurred is often the harder problem. This session uses PonniRAG (https://ponniarchive.com/), a production open-source RAG system built over a historical Tamil magazine archive from 1947–1955, containing 1,700+ articles across 108 issues and 487 contributors, to walk through the debugging process from query to answer.
The session examines how each stage of the RAG pipeline behaves differently when working with Tamil, a morphologically rich language with historical spelling variations and OCR-generated text. The session will focus on:
- Query understanding & routing: identifying retrieval questions versus metadata lookups and handling Tamil-specific question patterns.
- Filtering & retrieval: dealing with morphological variations, embedding-model conventions, recall and precision, and why hybrid search does not always solve the problem.
- Chunking & context assembly: handling article boundaries, OCR noise, historical context, conflicting information, and preserving metadata for provenance and citations.
- Evaluation & observability: identifying failures that are invisible at the model level and validating whether the evaluation tooling itself works correctly for non-Latin languages.
The session is grounded in real production failures and the fixes that were implemented, rather than a generic RAG architecture or a list of best practices. One example is an evaluation metric that reported near-zero scores because its tokenizer stripped non-Latin characters; another was the need to create a 50-question native-language ground-truth set because a suitable benchmark did not exist.
Attendees will leave with a practical, stage-by-stage debugging approach that can be applied to their own RAG systems, including:
- Localizing the failure before tuning the model
- Evaluating retrieval, chunk quality and context independently
- Treating embedding conventions and metadata as correctness requirements
- Auditing evaluation and observability tooling for multilingual systems
Although Tamil is the primary case study, the approach is applicable to practitioners building AI systems for other underserved Asian languages where tooling, benchmarks and language-specific infrastructure are limited.
Abinaya Mahendiran
CTO - Nunnari Labs| Independent Researcher | Community Enabler @ AI Tamil Nadu
Coimbatore, India
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top