Session
Building AI Inference Infrastructure for Kubernetes: Lessons Learned from a Maintainer’s Journey
AI inference is evolving at an incredible pace, and so are the requirements for running it efficiently on Kubernetes. Features that seemed sufficient a year ago quickly became limiting as new serving architectures, routing strategies, and inference engines emerged.
In this session, we'll cover the lessons we've learned while building Kubernetes-native AI inference infrastructure through the Kubernetes Inference Gateway (Gateway API Inference Extension) and llm-d.
We'll start with the initial architecture that addressed the immediate challenges. From there, we'll explore how we evolved it into a pluggable framework that allows new routing strategies to be supported without changing the core system. We'll conclude by exploring how this architecture continues to evolve to support emerging inference patterns, including Prefill/Decode disaggregation and beyond.
Whether you're building AI platforms, contributing to open source, or designing extensible Kubernetes systems, you'll leave with practical lessons on balancing simplicity, extensibility, and long-term evolution in a rapidly changing ecosystem.
Presented at KCD Porto 2026:
https://kcd-porto-2026.sessionize.com/session/1298879
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top