© Mapbox, © OpenStreetMap
Reza Jelveh

Reza Jelveh

Solutions Architect @ Dynamia.ai / HAMi

Taipei, Taiwan

Actions

Reza Jelveh is a solutions engineer at HAMi (CNCF Incubation), working on GPU workload mechanics in Kubernetes.

His background spans graphics drivers, kernel reverse engineering, satellite layer 2 reverse engineering, and infrastructure across semiconductor testing, healthcare, and data centers. He has been CTO for startups and public-sector infrastructure, navigating bare-metal through legacy middleware.

Area of Expertise

  • Finance & Banking
  • Information & Communications Technology
  • Manufacturing & Industrial Materials

Topics

  • Kubernetes
  • Edge
  • edge ai
  • Kubernetes Security
  • Container and Kubernetes security
  • Database
  • performance tuning
  • Cloud Native & Kubernetes

Shared GPU Scheduling & Proactive Autoscaling: A Production Blueprint for 1000+ GPUs

SNOW Corp. operates 1,000+ A100 GPUs serving 200 million users across three top-ranked GenAI applications (Snow, Epik, B612), handling 1,200+ AI workflows subject to extreme traffic volatility from viral AI trends.

The core bottleneck was Kubernetes' native GPU scheduling, which treats GPUs as atomic resources — forcing a 2x over-provisioning penalty on Train-to-Inference pipelines with no reliable visibility into actual GPU saturation.

This talk covers integrating HAMi for vGPU virtualization and extending KEDA with a custom Consumer Saturation metric for proactive autoscaling, hiding warm-up latency by scaling before traffic arrives.

We’ll detail the implementation: scheduler config, Prometheus metrics, and multi-region scaling via Helm GitOps. Results: >50% GPU waste cut and 91% faster recovery during surges. You’ll get a production blueprint for efficient, shared GPU platforms.

Vendor-Neutral GPU Observability for Kubernetes

GPU observability in Kubernetes is fragmented. Each vendor ships their own metrics stack: NVIDIA DCGM, AMD ROCm SMI. Platform teams running heterogeneous clusters stitch together three different dashboards to answer one question: "Are my GPUs being used efficiently?"

HAMi (CNCF Incubation) sits at the scheduling layer and sees every GPU operation. That single integration point gives you centralized observability across NVIDIA, AMD, Ascend, and any accelerator with a device plugin. Rather than scraping vendor-specific endpoints, HAMi instruments the scheduling path itself and reports utilization, memory pressure, and allocation efficiency per workload regardless of the underlying hardware.

This talk covers:
- Why GPU observability is harder than CPU observability (CUDA context model, MIG partitioning, device-plugin opacity)
- How HAMi's scheduling-layer instrumentation provides vendor-neutral metrics without per-vendor exporters

From Project to Production: HAMi and Viettel Cloud

Kubernetes treats GPUs as atomic resources, forcing over-provisioning and low utilization in multi-tenant AI Notebooks. DRA and HAMi's vGPU virtualization solve this, but only if implemented correctly.

Part 1: Mechanics of GPU Sharing. How DRA alters resource requests and HAMi implements fractional GPU allocation. The hardware constraints: memory isolation, compute slicing.

Part 2: Production at Viettel Cloud. Deployment architecture, bottlenecks moving from test to production, and operational realities of fractional GPUs for data science workloads at telco scale.

GPU Virtualization: First Setup and Understanding the Device Plugin / DRA

HAMi is a CNCF Incubation project for heterogeneous GPU virtualization in Kubernetes, supporting 10+ GPU/NPU vendors.

Attendees set up HAMi on their laptop - with or without a GPU - and trace how the device plugin intercepts GPU requests, allocates vGPU resources, and integrates with the scheduler.

We follow the HAMi architecture and GPU virtualization model (project-hami.io/docs/core-concepts/gpu-virtualization), then explore the DRA integration path. Laptop with Go, kind, and a GitHub account required.

GPU Virtualization: First Setup and Understanding the Device Plugin / DRA

HAMi is a CNCF Incubation project for heterogeneous GPU virtualization in Kubernetes, supporting 10+ GPU/NPU vendors.

Attendees set up HAMi on their laptop - with or without a GPU - and trace how the device plugin intercepts GPU requests, allocates vGPU resources, and integrates with the scheduler.

We follow the HAMi architecture and GPU virtualization model (project-hami.io/docs/core-concepts/gpu-virtualization), then explore the DRA integration path. Laptop with Go, kind, and a GitHub account required.

AI, Edge, and Storage Walk into a Mongolian Mine

Being able to interpret and mitigate seismic activity in mines can drastically improve safety for workers. However, mines are complex, noisy, and resource constrained environments, leading to suboptimal data. The computing environment can also be challenging with limited bandwidth and lack of modern computing equipment. This talk covers our journey in building a cloud native edge AI stream processing platform to analyze and interpret seismic activity in real time.

We will discuss overcoming the challenges of an industry heavily reliant on proprietary data formats and API’s, and of deploying Kubernetes (and other technologies) in air-gapped and low-resource environments, where cloud native storage goes right (and wrong). We will also demonstrate how we simulated our environment in the cloud and the benefits this brought to our deployment. The audience will walk away with a few nuggets of gold on how we created a real time decision making platform ready for the Gobi desert.

Project Lightning Talk + ContribFest + Maintainer Track: KubeCon + CloudNativeCon NA 2026 Sessionize Event Upcoming

November 2026 Salt Lake City, Utah, United States

KubeCon + CloudNativeCon North America 2026 Sessionize Event Upcoming

November 2026 Salt Lake City, Utah, United States

Observability Summit Europe 2026 Sessionize Event Upcoming

October 2026 Prague, Czechia

KubeCon + CloudNativeCon Japan 2026 Sessionize Event

July 2026 Yokohama, Japan

KCD & OpenInfra Days Vietnam 2026 Sessionize Event

July 2026 Hanoi, Vietnam

Reza Jelveh

Solutions Architect @ Dynamia.ai / HAMi

Taipei, Taiwan

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top