Akshay Pratinav
Intuit, Senior Staff Software Engineer
Mountain View, California, United States
Actions
I’m a Senior Staff Software Engineer with deep expertise in cloud‑native distributed computing. My career spans critical roles building internal developer productivity tools, where I’ve consistently delivered impactful, high‑performance solutions. In my current role at Intuit, I am leading the development of an internal Disaster Recovery platform that’s a non‑negotiable imperative for maintaining business continuity and customer trust. I also played a key role in the ITSM initiative, automating and streamlining incident management and significantly reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR) for major incidents. Additionally, I was instrumental in the design and development of the Policy Enforcement platform, providing a sophisticated, unified, and consistent experience for authoring, deploying, and enforcing policies across diverse cloud and Kubernetes infrastructure—embedding security, compliance, and operational efficiency. Outside of work, I enjoy travelling and watching movies.
Area of Expertise
Topics
EWOK: A Flexible Agent-Based Disaster Recovery Orchestration System for Stateful Cloud
When a regional incident strikes, every minute of downtime carries real customer impact — yet most organizations still rely on fragmented runbooks and manual coordination. This session presents EWOK, a production-grade, agent-based DR orchestration system built at Intuit to automate complex, stateful failovers across Active/Passive multi-region cloud architectures.
We'll walk through how EWOK replaces brittle DR scripts with declarative YAML workflows, an event-driven AWS State Machine, and a distributed Semaphore Service enforcing capacity-aware ordering across up to 100 concurrent multi-workload failovers — coordinating Kubernetes scaling, Aurora Global Database promotions, Route53 DNS updates, and Redis failovers while preventing GitOps configuration drift during live incidents.
Results from Intuit's critical financial platforms: MTTR reduced from 30+ minutes to under 5, zero idle hot-standby waste through just-in-time scaling, and 100% automated auditability. Attendees leave with a reusable architectural blueprint for production-hardened, platform-aware DR automation.
The Multi‑Workload Challenge: Predictable Failover Across the Stack
As Intuit’s financial platforms scale globally, disaster recovery and failover must work reliably across heterogeneous workloads—stateless services, stateful data tiers, scheduled jobs (cron), and asynchronous/background pipelines. This session shares how we translate RTO/RPO objectives into concrete multi‑region architectures (active‑active and active‑passive), global traffic steering, and data continuity strategies that preserve correctness under partial failures, network partitions, and cascading retries.
In this talk, I’ll share how we choose the right failover approach by workload and criticality, how we operationalize readiness with automation, health gates, and regular drills, and how we measure what matters so recovery is predictable, fast, and boring. You’ll leave with patterns that work, pitfalls to avoid, and a simple decision checklist you can adapt to steadily raise resilience in your own organization.
Self-Healing Systems: How LLM Agents Are Reinventing Cloud-Native Disaster Recovery
Cloud-native systems have outgrown the incident response playbook. Static runbooks, manual operator intervention, and reactive on-call rotations weren't designed for the scale and complexity of modern distributed systems — and the result is longer outages, inconsistent recoveries, and burned-out engineers.
This session introduces an agentic AI approach to disaster recovery, where LLM-based agents don't just advise — they act. Multiple collaborating agents work in concert to detect anomalies, reason about root causes, evaluate recovery strategies, and execute remediation directly through your existing Kubernetes and cloud interfaces. Human oversight remains intact through safety controls for high-impact operations, so you keep the guardrails without the bottlenecks.
We'll show real-world results: dramatic reductions in MTTR and operational toil, with measurable improvements in recovery consistency and reliability. More importantly, we'll walk through the architecture — how agents are orchestrated, how they reason under uncertainty, and how this system evolves from a passive advisory tool into an autonomous SRE co-pilot.
If you're an SRE, platform engineer, or architect wondering where AI fits into your reliability story, this session gives you a concrete, battle-tested answer.
OPEN Session: The Multi‑Workload Challenge: Predictable Failover Across the Stack
As Intuit’s financial platforms scale globally, disaster recovery and failover must work reliably across heterogeneous workloads—stateless services, stateful data tiers, scheduled jobs (cron), and asynchronous/background pipelines. This session shares how we translate RTO/RPO objectives into concrete multi‑region architectures (active‑active and active‑passive), global traffic steering, and data continuity strategies that preserve correctness under partial failures, network partitions, and cascading retries.
In this talk, I’ll share how we choose the right failover approach by workload and criticality, how we operationalize readiness with automation, health gates, and regular drills, and how we measure what matters so recovery is predictable, fast, and boring. You’ll leave with patterns that work, pitfalls to avoid, and a simple decision checklist you can adapt to steadily raise resilience in your own organization.
Faster Recovery, Less Toil: Rethinking DR with Agentic AI
Cloud-native infrastructure has outgrown the era of static runbooks and on-call engineers triaging alerts at 2am. As system complexity compounds, traditional disaster recovery falls further behind—leaving teams trapped in reactive loops, extended outages, and mounting on-call fatigue.
This talk introduces an agentic AI framework for disaster recovery, where LLM-based agents move beyond passive recommendations to actively participate in failure diagnosis and recovery. Rather than executing predefined scripts, a network of collaborating agents continuously monitors for anomalies, reasons over probable root causes, weighs remediation options, and drives recovery actions through native cloud and Kubernetes interfaces—all while preserving human oversight for high-blast-radius operations.
The results speak for themselves: faster recovery times, less operational toil, and more consistent incident response—even under pressure. More broadly, this work charts a path for how agentic AI transitions from advisory assistant to active SRE co-pilot, powering a new generation of autonomous, self-healing cloud operations.
If you're looking to move beyond reactive firefighting and toward incident workflows that are intelligent, auditable, and genuinely resilient—this session is for you.
Faster Recovery, Less Toil: Rethinking DR with Agentic AI
Cloud-native infrastructure has outgrown the era of static runbooks and on-call engineers triaging alerts at 2am. As system complexity compounds, traditional disaster recovery falls further behind—leaving teams trapped in reactive loops, extended outages, and mounting on-call fatigue.
This talk introduces an agentic AI framework for disaster recovery, where LLM-based agents move beyond passive recommendations to actively participate in failure diagnosis and recovery. Rather than executing predefined scripts, a network of collaborating agents continuously monitors for anomalies, reasons over probable root causes, weighs remediation options, and drives recovery actions through native cloud and Kubernetes interfaces—all while preserving human oversight for high-blast-radius operations.
The results speak for themselves: faster recovery times, less operational toil, and more consistent incident response—even under pressure. More broadly, this work charts a path for how agentic AI transitions from advisory assistant to active SRE co-pilot, powering a new generation of autonomous, self-healing cloud operations.
If you're looking to move beyond reactive firefighting and toward incident workflows that are intelligent, auditable, and genuinely resilient—this session is for you.
DevopsCon New York 2026 Upcoming
AI-Powered Service Operations: Intuit's Incident Management Story
KCD SF Bay Area 2026 Sessionize Event Upcoming
IEEE Cloud Summit 2026
EWOK: A Flexible Agent-Based Disaster Recovery Orchestration System for Stateful Cloud Services
IEEE SVCC 2026
One Engine to Rule Them All: How We Turned a Reliability Framework into a Security Powerhouse
AI Con 2026
From Reactive to Proactive: Intuit's Journey in AI-Powered Incident Management
PagerDuty on Tour 2026
From Reactive to Proactive: AI-Powered Incident Management at Enterprise Scale with Twilio and Intuit
SRE Day 2026
From Compute to Datastores: Predictable Region Failover for Microservice Stacks
All Things AI 2026
AI-Driven Disaster Recovery for Cloud-Native Systems
IEEE SoutheastCon 2026
AIDR-Cloud: An Agentic AI–Driven Incident Response Framework
DeveloperWeek 2026 Sessionize Event
Conf42 Devops 2026
The Multi‑Workload Challenge: Predictable Failover Across the Stack
API World 2025
Scalable Region Failover for Microservices: Platform-Level Resilience at Intuit
Elastic ECK Roundtable 2022
Adobe’s journey on adopting ECK stack for building realtime observability
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top