Session
Topology-Aware Scheduling for AI Training & Inference with Kueue
Kueue, the Kubernetes-native workload orchestrator, helps run AI workloads at scale with multi-tenant quota management and advanced scheduling.
In this session, we show through a case study at the Meta Superintelligence Lab how Kueue provided the right quota and scheduling layer for large-scale AI training and inference across complex, heterogeneous GPU topologies, enabling a shift from an in-house platform to an open-source stack on Kubernetes.
We then dive into Kueue’s multi-layer Topology-Aware Scheduling (TAS) for modern clusters with hierarchical network fabrics. We cover the main concepts and algorithms behind Hierarchical TAS, and show why multi-layer TAS was critical for this transition.
We close with other cutting-edge Kueue capabilities, including Workload-Aware Scheduler integration, elastic workloads, and multi-cluster job scheduling with MultiKueue.
Wei Huang
Ex-lead of Kubernetes sig-scheduling
Menlo Park, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top