Session
Your AI Workload Was Admitted. Why Didn't It Run?
A Kubernetes workload can be admitted and still fail to become useful compute. Across validation of Job, RayJob, PyTorchJob, JobSet, and AppWrapper, we found that apparently similar "stuck" workloads can come from very different layers: API and CRD readiness, controller reconciliation and RBAC, or Kueue admission and quota accounting.
This technical session organizes those implemented integration lessons into a three-gate diagnostic method. It explains how API and CRD readiness, controller reconciliation, and Kueue admission can be checked independently, and how the status returned at each layer narrows the failure domain. It also covers verified implementation findings: framework-specific placement behavior, an RBAC failure that prevented reconciliation, and a DRA quota-accounting gap where a workload could receive a GPU without consuming the expected Kueue quota.
Attendees leave with a practical test sequence and failure taxonomy for onboarding heterogeneous AI workload controllers without treating admission as proof that useful compute will run.
Karthik Ravi
Senior Software Engineer, AI/ML Infrastructure at PayPal
Mountain View, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top