Session
Bridging the Semantic Gap in Text-to-SQL Systems
Text-to-SQL technology promises to democratize data access, but its effectiveness has been largely confined to English. When deployed in multilingual environments, these systems suffer a dramatic drop in performance, not just in execution accuracy, but more critically, in semantic alignment, failing to capture the user's true intent. This session dives deep into a novel framework that bridges this multilingual gap, presenting a production-oriented approach that moves beyond brittle translation-based methods and static prompting.
At the core of this solution is a sophisticated reinforcement learning strategy (Group Relative Policy Optimization - GRPO) combined with a groundbreaking contrastive reward signal. This semantic reward, powered by a multilingual encoder, teaches the model to prioritize the meaning and intent of a query, regardless of the source language. We will explore how this focus on semantic fidelity, rather than simple execution accuracy, leads to the generation of robust, precise, and reliable SQL queries that hold up even when the underlying database schema changes.
The most compelling aspect of this approach is its efficiency. We will demonstrate how a small, 3B parameter model, fine-tuned with this framework on only 3,000 examples, significantly outperforms a much larger 8B zero-shot model. This isn't just about better accuracy; it's about achieving it with a fraction of the computational cost, making truly global Text-to-SQL systems both practical and affordable.
Ashish Kattamuri
Staff Software Engineer, Proofpoint
Denver, Colorado, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top