Session

Federated llm-d: Elevating distributed inference beyond cluster boundaries

As generative AI models trend towards becoming commodity, the ability to serve them in a cost-effective high-performant fashion across hardware accelerators impacts time to market for Enterprise AI Initiatives. Llm-d is a popular open-source AI Framework that meets this exact need.

Models are sensitive to hardware accelerators. Hardware accelerators are expensive and in short supply. Model serving time cannot and should not be impacted by unavailability of desired accelerator in a given Kubernetes cluster. The desired accelerator might just be a hop away in a different cluster in same/different region in same/different cloud. It would be prudent to source accelerators from all available clusters to deliver state-of-the-art performance to meet Enterprise SLAs.

This talk presents a federated llm-d framework which stretches llm-d across multiple clusters. We will discuss use cases for federated distributed inference, and present a blueprint for the federated stack along with a demo.

Abhishek Malvankar

Senior Software Engineer, Master Inventor at IBM Research

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top