Session

Route, Serve, Adapt, Repeat: Adaptive Routing for AI Inference Workloads in Kubernetes

Running inference on K8s can be costly and extremely slow.
Today’s inference routing strategies like traffic splitting, node affinity or session stickiness — are all static. Once defined, they ignore changing load, queue build-ups, and cache locality.

Inference workloads, however, are dynamic: requests vary, cache states shift, and cluster conditions evolve. Static routing strategies simply can’t keep up, leading to latency spikes and wasted GPU cycles.

With K8s Gateway API Inference Extension, we introduce adaptive routing strategies for inference, driven by real-time signals such as queue length and cache utilization. By continuously adapting, the system balances cache efficiency with load distribution, reduces latency, improves GPU utilization, and lowers costs at scale.

Attendees will learn why static routing strategies limit inference performance and see benchmarks demonstrating latency, efficiency, and cost gains with adaptive routing in K8s Gateway API Inference Extension.


Presented at KubeCon EU 2026, Amsterdam:
https://kccnceu2026.sched.com/event/2CW2C/route-serve-adapt-repeat-adaptive-routing-for-ai-inference-workloads-in-kubernetes-nir-rozenbaum-ibm-kellen-swain-google?iframe=yes&w=&sidebar=yes&bg=no

Nir Rozenbaum

Kubernetes & Open Source, NVIDIA

Yoqneam, Israel

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top