Session

“A Kubestronaut Riding a Dragon”: Scaling Generative AI on GKE

Generative AI is quickly becoming a production workload, but running models like Stable Diffusion XL on GKE introduces challenges around GPUs, scheduling, and performance.

In this session, we present a real-world implementation of an AI image generation platform running on Kubernetes. You’ll learn how to containerize, deploy, and scale inference workloads using open-source tools like Diffusers, PyTorch, and Streamlit.

We’ll explore GPU vs CPU performance, model loading behavior (~8GB), and cold start impacts, along with how Kubernetes distributes and scales inference workloads.

The talk includes a live demo generating images in real time (including “a kubestronaut riding a dragon in space”) and shares practical insights for running generative AI workloads on Kubernetes in any environment.

Yongkang He

Founder @KSUG.AI @KubeSmart.AI | Creator @awstronaut @kubestrong

Singapore

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top