Session
From old-school HPC to sbatch: regulated industries using immutable systems & confidential computing
Getting a GPU cluster from racked to running distributed training is where weeks disappear. A specialised AI-cluster recipe can carry 268 config values across 16 components, and a single mismatch can cost one or two percentage points of training throughput. We treat that recipe as an artifact you capture, version-lock, and replay, then prove it live. On stage, on a bare-metal H100 cluster, we bootstrap baseline Kubernetes, snapshot it with NVIDIA AI Cluster Runtime (AICR) into a version-locked recipe, render it as Helm values for Spectro Cloud's PaletteAI (a declarative Cluster API platform), deploy Slurm via the Slinky slurm-operator, and submit a multinode sbatch job, racked to running, end to end.
You leave with a workflow you can replay on any conformant Kubernetes cluster, validated against the CNCF Kubernetes AI Conformance requirements, and AI/HPC defaults you'd otherwise learn the hard way. For AI-infrastructure maintainers bridging Slurm-native HPC and cloud-native AI.
Eduardo Arango Gutierrez
Senior Systems Software Engineer @NVIDIA
Landsberg am Lech, Germany
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top