Session
Designing High-Performance RAG Systems on Kubernetes
Retrieval-Augmented Generation (RAG) is changing how we build AI products: instead of relying on a model’s “memory,” we ground responses in fresh, domain-specific knowledge. The demo is easy. Shipping it reliably at scale is not.
In this talk, we’ll break down how to design, build, and scale end-to-end RAG pipelines using cloud-native patterns and modern AI infrastructure. We’ll explore reference architectures for retrieval and ranking services, model serving strategies, and orchestration best practices, then zoom into the engineering tradeoffs that determine whether your system feels instant or (painfully) slow.
Expect practical guidance on keeping latency low and availability high: timeouts and backpressure, semantic caching, fallback strategies, and the observability stack you’ll want before users show up. The focus is framework, and language-agnostic, so you can apply these patterns whether you’re working in JVM, Python, Go, Node, or anything else.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top