Session

Scaling AI Inference Across Remote Clusters with VLLM

AI models are growing in size and complexity, making efficient workload distribution essential for scalable inference. This session explores how to deploy and manage AI inference workloads across remote Kubernetes clusters using VLLM, an optimized solution for high-throughput inference.

Attendees will learn:
- How VLLM improves AI inference performance
- Techniques for distributing workloads dynamically across remote clusters
- Strategies for balancing latency, compute availability, and cost

This talk is ideal for AI engineers, platform architects, and open-source enthusiasts looking to optimize AI inference across distributed infrastructure.

Andy Anderson

IBM Research, KubeStellar Community Maintainer

Stamford, Connecticut, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top