Session

Shared GPU Scheduling & Proactive Autoscaling: A Production Blueprint for 1000+ GPUs

SNOW Corp. operates 1,000+ A100 GPUs serving 200 million users across three top-ranked GenAI applications (Snow, Epik, B612), handling 1,200+ AI workflows subject to extreme traffic volatility from viral AI trends.

The core bottleneck was Kubernetes' native GPU scheduling, which treats GPUs as atomic resources — forcing a 2x over-provisioning penalty on Train-to-Inference pipelines with no reliable visibility into actual GPU saturation.

This talk covers integrating HAMi for vGPU virtualization and extending KEDA with a custom Consumer Saturation metric for proactive autoscaling, hiding warm-up latency by scaling before traffic arrives.

We’ll detail the implementation: scheduler config, Prometheus metrics, and multi-region scaling via Helm GitOps. Results: >50% GPU waste cut and 91% faster recovery during surges. You’ll get a production blueprint for efficient, shared GPU platforms.

Reza Jelveh

Solutions Architect @ Dynamia.ai / HAMi

Taipei, Taiwan

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top