Session
Interactive Is Not Batch: Scheduling Multi-Cluster JupyterHub with Kueue
Batch schedulers assume a submitted job can wait quietly. Interactive notebooks cannot: users need visible admission progress, bounded startup behavior, cancellation, cleanup, and reliable reconnects while capacity may live in another Kubernetes cluster. Treating a notebook Pod like an ordinary queued job creates failure modes at the seams between JupyterHub, Kueue, and multi-cluster placement.
This session builds an upstream-only control-plane model for Kueue-aware interactive sessions. We trace a spawn request through context selection, queue admission, workload progress, notebook startup, proxy routing, cancellation, and stale-event cleanup. We then separate four problems that are often conflated: admission state, user-visible progress, event-loop responsiveness, and cross-cluster lifecycle ownership.
Attendees leave with a lifecycle state machine, failure taxonomy, and synthetic test plan for validating interactive workloads across Kubernetes clusters—without relying on private topology or vendor-specific components.
Karthik Ravi
Senior Software Engineer, AI/ML Infrastructure at PayPal
Mountain View, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top