Session

⚡ Fast Restarts, Not Just Fast Starts: Accelerating Pod Recovery

Pod startup optimization is no longer enough; restart latency is now a core bottleneck for reliability and cost, especially in AI/ML and batch workloads. This session connects start-up and restart as one lifecycle problem. We begin with a practical taxonomy (container restart vs Pod recreation vs rollout/eviction), then break down where restart time is lost: re-scheduling, image pull, init re-execution, and control-plane churn.

The main case study is JobSet in-place restart, where prototype results show recovery improving from 2m10s to 10s at 5,000-node scale. We map this to new Kubernetes capabilities such as per-container restart policies/rules and restart-all-containers. Finally, we provide a production checklist across kubelet tuning, workload controller design, probes/signals, and observability so teams can safely reduce restart time without harming SLOs.

Baofa Fan

DaoCloud, IT Engineer

Shanghai, China

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top