Session

Beyond GPU Counts: Proving Kubernetes DRA Workloads Are Actually Schedulable

Requesting nvidia.com/gpu: 2 treats accelerators as interchangeable device counts. Modern AI clusters need a richer contract for partitioned, topology-sensitive, or driver-managed devices. Kubernetes Dynamic Resource Allocation (DRA) provides that contract, but installing a DRA driver does not prove that real workloads can be admitted, allocated, run, and cleaned up correctly.

This session presents an upstream-only validation method for the path from ResourceClaim to accelerator compute. We separate readiness into six observable gates: driver availability, claim binding, Kueue admission, concurrent allocation, real GPU execution, and workload-controller lifecycle. We apply the same contract to a Job, JobSet, RayJob, and PyTorchJob to expose failures that installation checks miss.

The key lesson is that allocation and queue accounting are different correctness problems. Attendees leave with a reusable compatibility matrix, failure taxonomy, and rollout sequence for Kubernetes-based AI infrastructure.

Karthik Ravi

Senior Software Engineer, AI/ML Infrastructure at PayPal

Mountain View, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top