Session
Scaling AI Inference Across Remote Clusters with VLLM
AI models are growing in size and complexity, making efficient workload distribution essential for scalable inference. This session explores how to deploy and manage AI inference workloads across remote Kubernetes clusters using VLLM, an optimized solution for high-throughput inference.
Attendees will learn:
- How VLLM improves AI inference performance
- Techniques for distributing workloads dynamically across remote clusters
- Strategies for balancing latency, compute availability, and cost
This talk is ideal for AI engineers, platform architects, and open-source enthusiasts looking to optimize AI inference across distributed infrastructure.
Andy Anderson
IBM Research, KubeStellar Community Maintainer
Stamford, Connecticut, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top