Session
Multi-Site Kubernetes for On-Prem GPU Workloads
Public cloud couldn’t meet our shop-floor constraints: real-time Vision AI requires immediate decisions within strict manufacturing cycles, and factory video feeds must stay on-prem for data sovereignty.
At Ford Otosan, we run a multi-site K8s cluster with a centralized control plane and remote GPU nodes distributed across plants over a private network. By moving inference to our on-prem edge datacenters, we achieve the low-latency performance required for real-time intervention.
In this talk, we explore:
Architecture: Managing multi-site K8s and handling site-link failures.
GPU Strategy: Scaling inference (Triton/vLLM) with HPA, and optimizing sharing via MIG/Time-slicing.
Workload Segregation: Separating strictly on-prem serving from cloud-eligible batch pipelines.
This case study provides a blueprint for leveraging open-source tools to solve data privacy and latency challenges at the Edge.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top