Session
Fixing the LLM Cold Start Problem with NVIDIA Dynamo Snapshot
Autoscaling an LLM workload sounds simple: demand increases, so you add more GPU workers. But between deciding to scale and gaining more serving capacity, a new worker may spend significant time initializing the runtime, loading model weights, preparing kernels, and allocating memory.
This talk explores how NVIDIA Dynamo Snapshot addresses this cold-start problem by checkpointing an initialized inference worker and restoring it when new capacity is needed. We’ll look at how CRIU handles the host-side process state. At the same time, cuda-checkpoint captures the GPU-side state, and how these mechanisms work together to checkpoint and restore a running GPU-accelerated inference process.
We’ll walk through the checkpoint and restore lifecycle, from freezing a running worker and serializing its state to restoring that state onto a GPU and resuming execution. We’ll also discuss the systems challenges around snapshot storage, workload quiescing, external state, and restoring workers across nodes.
Finally, we’ll examine the performance impact: reducing inference worker startup time from approximately 90 seconds to 9 seconds.
The broader lesson is that inference autoscaling is not only about deciding when to add replicas. It is also about reducing the time between allocating expensive compute and turning that compute into useful serving capacity.
Rasheedat Atinuke Jamiu
AI Infrastructure engineer
Kaduna, Nigeria
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top