Session

From Reactive SRE to Predictive Reliability: Building Self-Healing Cloud Systems

Modern cloud systems are operating at a scale where traditional reactive SRE practices are no longer sufficient. Teams are overwhelmed by alerts, delayed incident response, and increasing system complexity across distributed and multi-cloud environments.

In this session, I will share how we evolved reliability engineering from reactive troubleshooting into a predictive and automated system. Using real-world production examples, we will explore how to detect failures before they happen and build systems that can recover automatically.

We will cover practical approaches to:

Designing self-healing systems using automation and observability
Reducing alert fatigue with intelligent signal correlation
Automating incident response and recovery workflows
Handling large-scale reliability challenges in Kubernetes and cloud-native environments

This talk focuses on real implementation strategies, lessons learned, and patterns that can be applied immediately to improve system reliability and reduce operational overhead.

Charit Upadhyay

Adobe, Senior Site Reliability Engineer

San Francisco, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top