Session
From AI Prototype to Kubernetes: Production Guardrails for LLM Applications
AI applications can move from idea to working prototype quickly, but deploying them reliably introduces challenges that model demonstrations often hide. Real-time APIs, WebSocket connections, external LLM services, vector retrieval, authentication, background processing, and unpredictable workloads all create new operational requirements.
This session presents a cloud-native architecture for moving an API-driven generative AI application toward Kubernetes. Using a real-time AI assistant as the reference use case, the talk examines how to separate frontend, API, transcription, retrieval, orchestration, and persistence workloads into independently deployable components.
Attendees will learn how Kubernetes primitives can support autoscaling, secrets management, health checks, resource controls, rolling deployments, workload isolation, and failure recovery. The session will also cover challenges such as scaling stateful WebSocket sessions, managing external model rate limits, tracing requests across AI services, controlling inference costs, and designing graceful degradation when a model or transcription provider becomes unavailable.
The presentation is vendor-neutral and focuses on reusable Kubernetes and cloud-native patterns rather than a particular commercial platform.
Shiva Kalyan Reddy Giri
Fidelity Investments, Software Engineer
Dallas, Texas, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top