Session
Operating AI Workloads in Production: Reliability Lessons from Real Systems
Most AI talks focus on building models. Few discuss what happens after deployment.
As organizations move AI systems from experimentation to mission-critical production environments, new reliability challenges emerge: GPU resource contention, burst inference traffic, model rollout failures, cost volatility, observability blind spots, and security risks across data pipelines.
In this session, we’ll explore practical, real-world lessons from operating AI workloads at scale in cloud-native environments. Attendees will learn:
How to design resilient infrastructure for AI inference and training workloads
Kubernetes strategies for scaling GPU-based systems
Blue/green and canary deployments for models
Observability patterns for AI systems (latency, drift, cost monitoring)
Securing AI pipelines without slowing innovation
This talk bridges the gap between ML innovation and production reliability. Attendees will leave with actionable architectural patterns for running AI systems safely, securely, and at scale.
Charit Upadhyay
Adobe, Senior Site Reliability Engineer
San Francisco, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top