Session
Vendor-Neutral GPU Observability for Kubernetes
GPU observability in Kubernetes is fragmented. Each vendor ships their own metrics stack: NVIDIA DCGM, AMD ROCm SMI. Platform teams running heterogeneous clusters stitch together three different dashboards to answer one question: "Are my GPUs being used efficiently?"
HAMi (CNCF Incubation) sits at the scheduling layer and sees every GPU operation. That single integration point gives you centralized observability across NVIDIA, AMD, Ascend, and any accelerator with a device plugin. Rather than scraping vendor-specific endpoints, HAMi instruments the scheduling path itself and reports utilization, memory pressure, and allocation efficiency per workload regardless of the underlying hardware.
This talk covers:
- Why GPU observability is harder than CPU observability (CUDA context model, MIG partitioning, device-plugin opacity)
- How HAMi's scheduling-layer instrumentation provides vendor-neutral metrics without per-vendor exporters
Reza Jelveh
Solutions Architect @ Dynamia.ai / HAMi
Taipei, Taiwan
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top