Session

Voice Agents: Where Latency Hides and How to Fight It

A voice agent pipeline looks simple: speech-to-text, LLM, text-to-speech. In practice, most of the engineering effort goes into things that aren't on that diagram.
This talk breaks down the latency budget of a real-time voice agent - what has to happen between the user going silent and the agent starting to speak. I'll cover the parts that took us the longest to get right: detecting when someone actually stopped talking (harder than it sounds), streaming orchestration when components in your pipeline run at different speeds, handling interruptions without losing conversation context, and transport layer choices that affect how responsive the whole thing feels.
I'll go through specific patterns we landed on and be equally specific about what didn't work - approaches that looked solid in testing but fell apart with real users, and architectural choices we had to adapt. We'll also evaluate the newer speech-to-speech models and analyze their pros and cons for production deployments.
If you're a developer who might end up building anything voice-based in the next time, this is the context you'll wish you had going in.

Agata Chudzińska

CTO / AI Solutions Architect at theBlue.ai GmbH

Poznań, Poland

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top