Session
Beyond Model Serving: Running Production LLMs on Kubernetes with KServe
Running an LLM on Kubernetes is easy. Running it in production is a different story.
Once you move beyond a demo, you face challenges that go far beyond serving a model: GPU scheduling, traffic spikes, streaming responses, autoscaling, model updates, and keeping latency low while efficiently using expensive GPU resources. Production LLM workloads require new approaches to scaling, routing, resource management, and platform integration.
In this session, KServe maintainers will share how the project has evolved beyond traditional model serving. We'll discuss how inference engines, AI gateways, autoscaling, distributed serving, and traffic management work together through Kubernetes-native APIs to simplify running LLM workloads at scale.
We'll share lessons learned from building and operating these capabilities, including practical deployment patterns and design decisions from real production environment
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top