Session

Deploying Production-Ready, CNCF AI-Conformant LLM Inference Platforms on Kubernetes

Production LLM inference on Kubernetes requires more than running a model server. A reliable platform must coordinate healthy operators, schedulable GPUs, gateway routing, autoscaling behavior, accelerator metrics, and workload performance while remaining portable across managed and self-hosted Kubernetes clusters.

CNCF Kubernetes AI Conformance defines a baseline for running AI workloads reliably on Kubernetes. The program now supports layered products, allowing an inference platform built on a conformant Kubernetes distribution.

This session will show how NVIDIA’s open-source AI Cluster Runtime (AICR) turns a cluster snapshot into a recipe-driven, validated, and bundled deployment with overlays for production-ready LLM inference platforms.

We will provide an update on AI Conformance for layered platforms, then walk through deploying LLM inference platforms on K8s using AICR. We will show how AICR emits a structured evidence directory mapped to CNCF AI Conformance requirements.

Yuan Chen

Nvidia, Software Engineer, Kubernetes, GPU, AI/ML Infrastructure, Open Source

San Jose, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top