Session

Evolving Kubernetes Scheduling: Workload API, Gang Scheduling, and In-Place Resizes

As Kubernetes handles more complex batch and distributed ML workloads, the limits of default pod-by-pod scheduling become clear. Running these workloads efficiently often requires complex workarounds to manage resource contention.

In this session, Abdel examines how Kubernetes scheduling primitives are adapting. We will cover the Workload and PodGroup APIs, explaining how they enable gang scheduling, a practical requirement for distributed jobs that rely on "all-or-nothing" placement to prevent resource deadlocks.

We will also look at lifecycle management via In-Place Pod Resizing. We will demonstrate how to dynamically adjust CPU and memory for pods so they can run without triggering restarts, which is crucial for applications with slow initialization times.

Attendees will get a factual look at the current state of these features and how Google engineers evaluate them in large-scale environments.

Abdel Sghiouar

Cloud Developer Advocate

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top