Session
Prepare for Disruptions: How we upgrade the whole ML training fleet bi-weekly
Machine learning jobs are particularly vulnerable to node disruptions—whether planned (like host maintenance, kernel upgrades, or security patches) or unplanned (such as GPU ECC memory errors or sudden node failures). These interruptions can derail progress and waste valuable training time.
In this talk, we’ll explore how to build disruption-tolerant ML infrastructure on Kubernetes that balances platform reliability with job continuity. We’ll cover techniques we’ve developed and battle-tested at scale, including:
• Automatic Multi-stage checkpoint and restore of training jobs to allow fast and seamless recovery after interruptions.
• Intelligent scheduling and smart collocation to account for node health, job characteristics, and maintenance timing.
• Job-aware backpressure mechanisms that coordinate updates and reduce the likelihood of disruption during critical job phases.
Attendees will leave with practical strategies for managing infrastructure disruptions leveraging Kubernetes.
Ankit Goyal
LinkedIn, AI Platform, Principal Staff Software Engineer
Sunnyvale, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top