Session
Deploying Production-Ready, CNCF AI-Conformant LLM Inference Platforms on Kubernetes
Production LLM inference on Kubernetes requires more than running a model server. A reliable platform must coordinate healthy operators, schedulable GPUs, gateway routing, autoscaling behavior, accelerator metrics, and workload performance while remaining portable across managed and self-hosted Kubernetes clusters.
CNCF Kubernetes AI Conformance defines a baseline for running AI workloads reliably on Kubernetes. The program now supports layered products, allowing an inference platform built on a conformant Kubernetes distribution.
This session will show how NVIDIA’s open-source AI Cluster Runtime (AICR) turns a cluster snapshot into a recipe-driven, validated, and bundled deployment with overlays for production-ready LLM inference platforms.
We will provide an update on AI Conformance for layered platforms, then walk through deploying LLM inference platforms on K8s using AICR. We will show how AICR emits a structured evidence directory mapped to CNCF AI Conformance requirements.
Yuan Chen
Nvidia, Software Engineer, Kubernetes, GPU, AI/ML Infrastructure, Open Source
San Jose, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top