Session

Fixing GPU Time-Slicing on Kubernetes: Dynamic, Memory-Safe Sharing for AI Inference

Time-slicing is often the only viable way to share GPUs on commodity hardware or in cloud environments where MIG and MPS are unavailable. In practice, Kubernetes time-slicing relies on static configuration and does not constrain GPU memory, leading to noisy neighbors, unpredictable performance, and OOM failures in inference workloads.

This session shows how to close that gap by combining memory-aware controls with smarter pod placement. Using KAI for GPU-aware scheduling and HAMi for per-workload memory limit enforcement, we demonstrate how to prevent memory overcommit, stabilize performance, and safely share GPUs across inference workloads. Under the hood, Kyverno mutates pod specs to inject KAI and HAMi configuration on inference pods that should share GPUs.

This talk shows how to get safe, OOM-free GPU sharing on standard Kubernetes clusters without hardware upgrades. Just dynamic scheduling and enforced memory limits that work with the GPUs you already have.

Julien Semaan

VP R&D @Kubex | CNCF TAG DevEx Tech Lead

Montréal, Canada

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top