Session

Who's Governing the GPU Lane? Policy-as-Code for AI Inference Traffic in Kubernetes

Platform teams running AI on Kubernetes face a governance gap most tooling ignores. GPU-backed models get traffic they shouldn't. Tenant isolation breaks silently. Failover bypasses policy. Cost attribution doesn't exist. Nobody notices until the bill arrives, or the wrong model serves the wrong customer.

The culprit is the inference infrastructure layer: Gateway API and Inference Extension resources that route AI traffic, control which models are reachable, and isolate tenants.

This session shows platform engineers how to close the gap with Kubernetes-native policy-as-code. A technical demo walks through reusable admission-time patterns that transfer to any Gateway API implementation and policy engine.

Attendees leave with reusable governance patterns: tenant isolation guardrails for InferencePool and InferenceObjective, safe failover and model access controls, cost-aware GPU routing, and defense in depth that combines admission-time policy with runtime controllers.

Cortney Nickerson

Head of Community at Nirmata

Donostia / San Sebastián, Spain

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top