Session
Beyond GPU Counts: Proving Kubernetes DRA Workloads Are Actually Schedulable
Requesting nvidia.com/gpu: 2 treats accelerators as interchangeable device counts. Modern AI clusters need a richer contract for partitioned, topology-sensitive, or driver-managed devices. Kubernetes Dynamic Resource Allocation (DRA) provides that contract, but installing a DRA driver does not prove that real workloads can be admitted, allocated, run, and cleaned up correctly.
This session presents an upstream-only validation method for the path from ResourceClaim to accelerator compute. We separate readiness into six observable gates: driver availability, claim binding, Kueue admission, concurrent allocation, real GPU execution, and workload-controller lifecycle. We apply the same contract to a Job, JobSet, RayJob, and PyTorchJob to expose failures that installation checks miss.
The key lesson is that allocation and queue accounting are different correctness problems. Attendees leave with a reusable compatibility matrix, failure taxonomy, and rollout sequence for Kubernetes-based AI infrastructure.
Karthik Ravi
Senior Software Engineer, AI/ML Infrastructure at PayPal
Mountain View, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top