Session

Accelerate your AI/ML Workloads with Topology-Aware Scheduling in Kueue

Optimizing execution time of AI training and inference is crucial in the era of LLMs. The workloads often exchange huge amounts of data between pods, making the network throughput a bottleneck.

Data centers have hierarchical organization with multiple layers, such as racks or blocks, however, leveraging this fact in vanilla Kubernetes is challenging as the scheduler needs to be aware of both workloads and the cluster topology. Kueue, as a Job-level scheduler, is already workload-aware. To tackle the second challenge, we propose a convention for labeling nodes by cloud-providers or cluster administrators. Leveraging this information, Kueue optimizes Pod placement within a cluster, ordering Pods by indices to enhance the performance of AI frameworks using NCCL.

In this session, we introduce the key concepts and machinery behind Topology-Aware Scheduling (TAS) in Kueue. We also compare TAS with alternatives and present results on how using it improves execution time of AI workloads.

Michał Woźniak

Google, Software Engineer

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top