Session
How to manage xPU automatically in large scale
With the development of the AI, more and more applications, e.g. LLM, asks for a better solution of network and storage to work together with GPU. xPU, including both DPU and IPU, is one of well-known solution which provides offload, hardware acceleration and so on. But an AI server usually includes serveral devices to maximize the performance, e.g. 8 DPU + 10 xPU, which introduces the challenge to manage all of those devices. Thanks to kubernetes, the GPU ares managed will by device plugin and operator; but xPU management is more complex which Including provisioning, storage and network.
This session will demonstrate how to manage those xPU in a large scale by some use cases and share the pain-point and solution accordingly; this session will also include part of the roadmap of the system, and general tips for xPU.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top