Session

The Grounding Verifier: How We Stopped Our AI From Lying About Money

We built an AI research tool that quotes specific numbers — stock prices, P/E ratios, filing dates, insider transaction sizes — back to retail investors. Every wrong number is a legal problem, a reputational problem, or both. After watching off-the-shelf LLMs confidently hallucinate revenue figures from 10-K filings, we built a grounding verifier that audits structured claims against the underlying tool output before the response leaves the server. Unverifiable claims get flagged and surfaced to the user, never silently passed.

This talk is the production walkthrough at the level of architecture, not implementation. The three failed approaches we tried first (regex pre-filters, LLM-as-judge audits, strict schema validation) and why each broke in instructive ways. The fail-closed instinct that took six weeks to internalize. The principled scope of structured verification — what kinds of claims a per-claim verifier can defend, and what kinds it structurally cannot. The specific failure category we caught early, including the time the model invented an entire insider sale that never happened.

By the end you'll know how to think about building a verifier for your own production agent, what the architectural tradeoffs are, and the specific bug shapes that will bite you in week three. If you've ever shipped an AI feature and then prayed nobody quoted it back to a regulator, this talk is for you.

Yash Shah

Founder & MCP Developer

New City, New York, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top