Session

GitOps for AI Infrastructure: The Operational Gaps Nobody Talks About

Large language models fit naturally into Kubernetes demos. Operating them reliably is a different problem.

Most operational challenges AI workloads introduce have individual mitigations: model caching, GPU autoscaling, persistent storage, artifact replication. The operational gap is how they compound. Solving them individually still leaves platform teams with unreliable deployments.

This poster examines the assumptions GitOps and Kubernetes tooling quietly depend on that AI inference workloads break: artifacts are small enough for standard operations, pods start fast enough for default probes, sync completion implies readiness, infrastructure state is disposable between deploys, and dependencies are visible and explicit.

Drawing from operating multi-cluster AI inference infrastructure across multiple environments, this poster presents deployment lifecycle diagrams, failure-mode taxonomies, and practical operational patterns for improving AI infrastructure reliability.

Kim Schaefer

Senior DevOps Engineer

Austin, Texas, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top