Michał Woźniak

Michał Woźniak

Google, Software Engineer

Actions

Michał is a software engineer with background in computer science, a PhD in computational biology, and 5+ years of professional experience. In his current role he is focusing on enhancing the support for batch workloads in the Kubernetes ecosystem. Outside of work he enjoys playing chess.

Accelerate your AI/ML Workloads with Topology-Aware Scheduling in Kueue

Optimizing execution time of AI training and inference is crucial in the era of LLMs. The workloads often exchange huge amounts of data between pods, making the network throughput a bottleneck.

Data centers have hierarchical organization with multiple layers, such as racks or blocks, however, leveraging this fact in vanilla Kubernetes is challenging as the scheduler needs to be aware of both workloads and the cluster topology. Kueue, as a Job-level scheduler, is already workload-aware. To tackle the second challenge, we propose a convention for labeling nodes by cloud-providers or cluster administrators. Leveraging this information, Kueue optimizes Pod placement within a cluster, ordering Pods by indices to enhance the performance of AI frameworks using NCCL.

In this session, we introduce the key concepts and machinery behind Topology-Aware Scheduling (TAS) in Kueue. We also compare TAS with alternatives and present results on how using it improves execution time of AI workloads.

WG-Batch Updates: What’s New and What Is Next?

Marcin and Michał will present improvements that the WG Batch has promoted in Kubernetes, and the opportunities under discussion to better support batch workloads such as HPC, AI/ML, data-analytics, etc.

We will discuss enhancements and improvements to the Job and JobSet APIs as well as new release and roadmap for Kueue, a Kubernetes subproject that offers job queueing and scheduling, to build a multitenant, multicluster batch system. Moreover we will update you on the progress of the Dynamic Resource Allocation (DRA) effort.

The WG Batch was created in 2022 to serve the demand from the ecosystem to better support batch applications in Kubernetes. The WG is composed of SIGs’ experts and developers from various communities, with the objective to set roadmaps and collaborate in designs and implementations.

Advanced resource management for running AI/ML workloads with Kueue

Kueue is a Job-level queueing manager which stands up to the challenges of managing computational resources to run batch workloads on Kubernetes.

We walk you through its architecture, demonstrating how it can be used to set up quota- and priority-based sharing of resources between multiple teams. We describe how the Kueue’s scheduler decides when to start or stop (preempt) a job.

We showcase Kueue by its production use at CyberAgent, where it is a building block of the multi-tenant system, supporting multiple engineers and ML research teams; using multiple types of CPUs and GPUs. Here, Kueue manages various types of Jobs (batch Job, MPIJob, or in-house Jobs), using various ML frameworks (TensorFlow, PyTorch or DeepSpeed).

Finally, we discuss the challenge of running ML training jobs which require all pods to be scheduled. We show how it is solved by using Kueue at CyberAgent, and how it can be solved using Kueue in the autoscaling environments with the new ProvisioningRequest API.

Michał Woźniak

Google, Software Engineer

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top