Session
Sovereign AI Inference on Kubernetes: From Zero-Touch Boot to Production LLM Serving
AI inference at scale often forces a trade-off between cloud convenience and data sovereignty. For regulated industries, this is non-negotiable: workloads must run in specific jurisdictions on verifiable infrastructure in production.
This session presents an open-source infrastructure stack, Kommodity, that addresses this directly. We show how Talos Linux and Cluster API form an immutable, API-driven GPU cluster foundation with network-based disk encryption and zero-touch HA bootstrap on private networks. On these clusters, we deploy Envoy, KServe and vLLM for model serving, MIG partitioning and bin-packing for GPU density, local model caching for fast cold starts, as well as the Gateway API Inference Extension for prefix-cache-aware routing.
Attendees will learn how to combine these projects into a production-grade inference platform, schedule GPU workloads across heterogeneous SKUs, and integrate zero-touch boot with inference operations, all without cloud vendor lock-in.
Steffen Karlsson
Principal Platform Engineer - Corti.ai
Copenhagen, Denmark
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top