Session

How LLMs Drift: Citation Observability for ChatGPT, Perplexity, Google AI Overviews, and Copilot

Most teams treat LLM output as deterministic. Run the same prompt twice, get roughly the same answer, ship it.

That assumption breaks the moment you start measuring. We've been running production-scale citation tracking across ChatGPT, Perplexity, Google AI Overviews, and Microsoft Copilot for 18 months, watching how the same prompts get different answers across engines, across time, and across model updates that nobody announced. The behavior is messier than the AI engineering community talks about publicly.

This session is what we found. It's not a marketing talk and it's not an evaluation framework pitch. It's the engineering reality of running an observability layer over four LLM systems that change underneath you, why our naive first architecture failed, and what enterprise customers actually need when they ask "am I being cited correctly?"

I'll cover the pipeline architecture (query distribution, citation extraction, storage, drift detection), the surprising patterns we've seen - including "Co-pilot's citation set for cybersecurity queries shifting 30%+ within a 6-hour window after a news event" and why the 11% domain overlap between ChatGPT and Perplexity citations forces a rethink of how you measure AI search visibility.

Useful if you're building eval infrastructure, customer-facing LLM analytics, RAG systems where citation matters, or anything where "did the model say what we expected?" needs to be answered for customers.

Deepak Gupta

Co-founder/CEO of GrackerAI

San Francisco, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top