Session

Operating AI Workloads in Production: Reliability Lessons from Real Systems

Most AI talks focus on building models. Few discuss what happens after deployment.

As organizations move AI systems from experimentation to mission-critical production environments, new reliability challenges emerge: GPU resource contention, burst inference traffic, model rollout failures, cost volatility, observability blind spots, and security risks across data pipelines.

In this session, we’ll explore practical, real-world lessons from operating AI workloads at scale in cloud-native environments. Attendees will learn:

How to design resilient infrastructure for AI inference and training workloads

Kubernetes strategies for scaling GPU-based systems

Blue/green and canary deployments for models

Observability patterns for AI systems (latency, drift, cost monitoring)

Securing AI pipelines without slowing innovation

This talk bridges the gap between ML innovation and production reliability. Attendees will leave with actionable architectural patterns for running AI systems safely, securely, and at scale.

Charit Upadhyay

Adobe, Senior Site Reliability Engineer

San Francisco, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top