Session
Scaling LLM Inference for Production Traffic
Most teams start with `vllm serve model-name`, but is that actually enough for production when the traffic burst is unpredictable, the GPU fills up with the KV cache, and 100s of users are wondering why it is taking so long to get a response?
This talk takes you on a ride from a basic vLLM server to a scalable inference stack. Rather than a feature showcase, we'd rather take a problem at every stage, apply optimizations, and benchmark it to check if it actually helps.
Finally, to make it scalable and production-ready, we'd add KEDA to scale replicas on meaningful demand signals, not just GPU usage. Lastly, we’d talk about the harsh reality of inference: model-loading cold starts, warm capacity, and safe scale-down.
Pratik Parmar
Developer Relations Professional who's passionate about helping developer communities across the globe!
Bengaluru, India
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top