Session
How We Run 8,000+ GPUs on Kubernetes with Slurm
As AI workloads scale to thousands of GPUs, organizations need scheduling capabilities that Kubernetes doesn't provide out of the box: gang scheduling for multi-node jobs, fair-share policies across teams, and topology-aware placement for optimal GPU communication. Slurm has solved these problems for decades in HPC, but traditionally requires dedicated bare-metal infrastructure.
In this talk, we show how to run Slurm natively on Kubernetes using Slinky, an open-source operator. This approach lets you keep existing Slurm workflows (job scripts, accounting, fair-share policies) while gaining Kubernetes benefits: declarative configuration, rolling updates, Helm deployments, and ecosystem integration. We operate this at scale with 8,000+ GPUs for distributed AI training.
Fagani Hajizada
Senior Software Engineer @ NVIDIA
Nürnberg, Germany
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top