Session
Slurm Meets Kubernetes: A Practical Guide to Cloud-Native HPC
Slurm powers most of the world's top supercomputers with capabilities beyond native Kubernetes: gang scheduling, topology-aware placement, fair-share policies, and job accounting. As AI scales up, these ecosystems can no longer operate in silos — Kubernetes gains battle-tested HPC scheduling, the HPC community gains cloud-native operations. It's time they converge.
We share the journey of bringing Slurm to Kubernetes through Slinky, an open-source operator that runs Slurm natively as Kubernetes workloads. I cover the obstacles — keeping state in sync across systems, handling node failures, enabling high-bandwidth GPU communication — and how we solved them. Whether you run four GPUs or four thousand, the patterns apply.
Attendees will learn why HPC scheduling and Kubernetes are stronger together, how to deploy a Slurm-enabled cluster with open-source Helm charts, and lessons from production. Whether you're a platform engineer, cluster admin, or just curious, this session is for you.
Fagani Hajizada
Senior Software Engineer @ NVIDIA
Nürnberg, Germany
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top