Session
Advanced resource management for running AI/ML workloads with Kueue
Kueue is a Job-level queueing manager which stands up to the challenges of managing computational resources to run batch workloads on Kubernetes.
We walk you through its architecture, demonstrating how it can be used to set up quota- and priority-based sharing of resources between multiple teams. We describe how the Kueue’s scheduler decides when to start or stop (preempt) a job.
We showcase Kueue by its production use at CyberAgent, where it is a building block of the multi-tenant system, supporting multiple engineers and ML research teams; using multiple types of CPUs and GPUs. Here, Kueue manages various types of Jobs (batch Job, MPIJob, or in-house Jobs), using various ML frameworks (TensorFlow, PyTorch or DeepSpeed).
Finally, we discuss the challenge of running ML training jobs which require all pods to be scheduled. We show how it is solved by using Kueue at CyberAgent, and how it can be solved using Kueue in the autoscaling environments with the new ProvisioningRequest API.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top