Nir Rozenbaum
Kubernetes & Open Source, NVIDIA
Yoqneam, Israel
Actions
Nir is an engineering leader with deep expertise in Kubernetes and Cloud-Native technologies.
He is a Maintainer of Kubernetes Inference Gateway, Chair of Kubernetes AI Gateway Working Group and SIG Lead & Maintainer of llm-d.
With a proven track record in CNCF open source, he drives innovation in AI workloads on Kubernetes and helps shape industry standards for Cloud-Native inference.
Links
Area of Expertise
Topics
Building AI Inference Infrastructure for Kubernetes: Lessons Learned from a Maintainer’s Journey
AI inference is evolving at an incredible pace, and so are the requirements for running it efficiently on Kubernetes. Features that seemed sufficient a year ago quickly became limiting as new serving architectures, routing strategies, and inference engines emerged.
In this session, we'll cover the lessons we've learned while building Kubernetes-native AI inference infrastructure through the Kubernetes Inference Gateway (Gateway API Inference Extension) and llm-d.
We'll start with the initial architecture that addressed the immediate challenges. From there, we'll explore how we evolved it into a pluggable framework that allows new routing strategies to be supported without changing the core system. We'll conclude by exploring how this architecture continues to evolve to support emerging inference patterns, including Prefill/Decode disaggregation and beyond.
Whether you're building AI platforms, contributing to open source, or designing extensible Kubernetes systems, you'll leave with practical lessons on balancing simplicity, extensibility, and long-term evolution in a rapidly changing ecosystem.
Presented at KCD Porto 2026:
https://kcd-porto-2026.sessionize.com/session/1298879
New Traffic, New Rules: Standardizing AI-Aware Networking in Kubernetes
Since launching the AI Gateway Working Group, the Kubernetes community has been turning early ideas about AI-aware networking into concrete designs and emerging API directions.
AI traffic challenges traditional networking assumptions. Requests now include prompts, tool calls, streaming responses, and external model interactions, requiring new patterns for routing, payload handling, and policy enforcement.
In this session, we will present the latest work from the AI Gateway WG, covering emerging AI traffic patterns, payload processing and transformation hooks, external model egress, and how these map to Gateway API and its extension model.
We will also discuss key design tradeoffs: what should be standardized vs left to implementations, how much payload awareness belongs in the network layer, and how to avoid premature abstractions.
Attendees will leave with a clear understanding of the WG's direction, current designs, and the open questions shaping AI networking in Kubernetes.
Presented at KubeCon NA 2026, Salt Lake City:
https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/program/schedule/?id=1289716
Route, Serve, Adapt, Repeat: Adaptive Routing for AI Inference Workloads in Kubernetes
Running inference on K8s can be costly and extremely slow.
Today’s inference routing strategies like traffic splitting, node affinity or session stickiness — are all static. Once defined, they ignore changing load, queue build-ups, and cache locality.
Inference workloads, however, are dynamic: requests vary, cache states shift, and cluster conditions evolve. Static routing strategies simply can’t keep up, leading to latency spikes and wasted GPU cycles.
With K8s Gateway API Inference Extension, we introduce adaptive routing strategies for inference, driven by real-time signals such as queue length and cache utilization. By continuously adapting, the system balances cache efficiency with load distribution, reduces latency, improves GPU utilization, and lowers costs at scale.
Attendees will learn why static routing strategies limit inference performance and see benchmarks demonstrating latency, efficiency, and cost gains with adaptive routing in K8s Gateway API Inference Extension.
Presented at KubeCon EU 2026, Amsterdam:
https://kccnceu2026.sched.com/event/2CW2C/route-serve-adapt-repeat-adaptive-routing-for-ai-inference-workloads-in-kubernetes-nir-rozenbaum-ibm-kellen-swain-google?iframe=yes&w=&sidebar=yes&bg=no
AI'm at the Gate! Introducing the AI Gateway Working Group in Kubernetes
Kubernetes Working Groups (WGs) play a vital role in shaping the future of Kubernetes and CNCF.
We’re excited to introduce a new addition: The AI Gateway WG.
This session will present the mission, scope, and early initiatives of the AI Gateway WG, focused on defining and advancing practices and standards at the intersection of AI and networking.
As AI systems increasingly rely on gateways, load balancers, and proxies, the WG is exploring the standardization of key capabilities, such as callouts to external AI backends for egress scenarios and payload-processing hooks that handle requests/responses before or after reaching an AI service. These primitives enable higher-level behaviors such as request/response guards, semantic routing, and other AI-aware traffic controls.
We’ll share the initial designs, early prototypes, and emerging directions shaping the WG’s roadmap.
Join us to learn how you can contribute and help shape the future of AI-aware gateway capabilities in Kubernetes.
Presented at KubeCon EU 2026, Amsterdam:
https://kccnceu2026.sched.com/event/2EF5t/aim-at-the-gate-introducing-the-ai-gateway-working-group-in-kubernetes-nir-rozenbaum-ibm-kellen-swain-google-morgan-foster-red-hat?iframe=yes&w=&sidebar=yes&bg=no
KCD Porto x DevOps Days Portugal 2026 Sessionize Event Upcoming
KubeCon + CloudNativeCon North America 2026 Sessionize Event Upcoming
KubeCon + CloudNativeCon Europe 2026 Sessionize Event
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top