© Mapbox, © OpenStreetMap
Akshay Pratinav

Akshay Pratinav

Intuit, Senior Staff Software Engineer

Mountain View, California, United States

Actions

I’m a Senior Staff Software Engineer with deep expertise in cloud‑native distributed computing. My career spans critical roles building internal developer productivity tools, where I’ve consistently delivered impactful, high‑performance solutions. In my current role at Intuit, I am leading the development of an internal Disaster Recovery platform that’s a non‑negotiable imperative for maintaining business continuity and customer trust. I also played a key role in the ITSM initiative, automating and streamlining incident management and significantly reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR) for major incidents. Additionally, I was instrumental in the design and development of the Policy Enforcement platform, providing a sophisticated, unified, and consistent experience for authoring, deploying, and enforcing policies across diverse cloud and Kubernetes infrastructure—embedding security, compliance, and operational efficiency. Outside of work, I enjoy travelling and watching movies.

Area of Expertise

  • Information & Communications Technology

Topics

  • Cloud & DevOps
  • SRE & AIOps
  • AI SRE
  • Incident Management
  • Change Management
  • AI-Driven Incident Management
  • Disaster Recovery

EWOK: A Flexible Agent-Based Disaster Recovery Orchestration System for Stateful Cloud

When a regional incident strikes, every minute of downtime carries real customer impact — yet most organizations still rely on fragmented runbooks and manual coordination. This session presents EWOK, a production-grade, agent-based DR orchestration system built at Intuit to automate complex, stateful failovers across Active/Passive multi-region cloud architectures.

We'll walk through how EWOK replaces brittle DR scripts with declarative YAML workflows, an event-driven AWS State Machine, and a distributed Semaphore Service enforcing capacity-aware ordering across up to 100 concurrent multi-workload failovers — coordinating Kubernetes scaling, Aurora Global Database promotions, Route53 DNS updates, and Redis failovers while preventing GitOps configuration drift during live incidents.

Results from Intuit's critical financial platforms: MTTR reduced from 30+ minutes to under 5, zero idle hot-standby waste through just-in-time scaling, and 100% automated auditability. Attendees leave with a reusable architectural blueprint for production-hardened, platform-aware DR automation.

The Multi‑Workload Challenge: Predictable Failover Across the Stack

As Intuit’s financial platforms scale globally, disaster recovery and failover must work reliably across heterogeneous workloads—stateless services, stateful data tiers, scheduled jobs (cron), and asynchronous/background pipelines. This session shares how we translate RTO/RPO objectives into concrete multi‑region architectures (active‑active and active‑passive), global traffic steering, and data continuity strategies that preserve correctness under partial failures, network partitions, and cascading retries.

In this talk, I’ll share how we choose the right failover approach by workload and criticality, how we operationalize readiness with automation, health gates, and regular drills, and how we measure what matters so recovery is predictable, fast, and boring. You’ll leave with patterns that work, pitfalls to avoid, and a simple decision checklist you can adapt to steadily raise resilience in your own organization.

Self-Healing Systems: How LLM Agents Are Reinventing Cloud-Native Disaster Recovery

Cloud-native systems have outgrown the incident response playbook. Static runbooks, manual operator intervention, and reactive on-call rotations weren't designed for the scale and complexity of modern distributed systems — and the result is longer outages, inconsistent recoveries, and burned-out engineers.

This session introduces an agentic AI approach to disaster recovery, where LLM-based agents don't just advise — they act. Multiple collaborating agents work in concert to detect anomalies, reason about root causes, evaluate recovery strategies, and execute remediation directly through your existing Kubernetes and cloud interfaces. Human oversight remains intact through safety controls for high-impact operations, so you keep the guardrails without the bottlenecks.

We'll show real-world results: dramatic reductions in MTTR and operational toil, with measurable improvements in recovery consistency and reliability. More importantly, we'll walk through the architecture — how agents are orchestrated, how they reason under uncertainty, and how this system evolves from a passive advisory tool into an autonomous SRE co-pilot.
If you're an SRE, platform engineer, or architect wondering where AI fits into your reliability story, this session gives you a concrete, battle-tested answer.

OPEN Session: The Multi‑Workload Challenge: Predictable Failover Across the Stack

As Intuit’s financial platforms scale globally, disaster recovery and failover must work reliably across heterogeneous workloads—stateless services, stateful data tiers, scheduled jobs (cron), and asynchronous/background pipelines. This session shares how we translate RTO/RPO objectives into concrete multi‑region architectures (active‑active and active‑passive), global traffic steering, and data continuity strategies that preserve correctness under partial failures, network partitions, and cascading retries.

In this talk, I’ll share how we choose the right failover approach by workload and criticality, how we operationalize readiness with automation, health gates, and regular drills, and how we measure what matters so recovery is predictable, fast, and boring. You’ll leave with patterns that work, pitfalls to avoid, and a simple decision checklist you can adapt to steadily raise resilience in your own organization.

Faster Recovery, Less Toil: Rethinking DR with Agentic AI

Cloud-native infrastructure has outgrown the era of static runbooks and on-call engineers triaging alerts at 2am. As system complexity compounds, traditional disaster recovery falls further behind—leaving teams trapped in reactive loops, extended outages, and mounting on-call fatigue.

This talk introduces an agentic AI framework for disaster recovery, where LLM-based agents move beyond passive recommendations to actively participate in failure diagnosis and recovery. Rather than executing predefined scripts, a network of collaborating agents continuously monitors for anomalies, reasons over probable root causes, weighs remediation options, and drives recovery actions through native cloud and Kubernetes interfaces—all while preserving human oversight for high-blast-radius operations.

The results speak for themselves: faster recovery times, less operational toil, and more consistent incident response—even under pressure. More broadly, this work charts a path for how agentic AI transitions from advisory assistant to active SRE co-pilot, powering a new generation of autonomous, self-healing cloud operations.

If you're looking to move beyond reactive firefighting and toward incident workflows that are intelligent, auditable, and genuinely resilient—this session is for you.

Faster Recovery, Less Toil: Rethinking DR with Agentic AI

Cloud-native infrastructure has outgrown the era of static runbooks and on-call engineers triaging alerts at 2am. As system complexity compounds, traditional disaster recovery falls further behind—leaving teams trapped in reactive loops, extended outages, and mounting on-call fatigue.

This talk introduces an agentic AI framework for disaster recovery, where LLM-based agents move beyond passive recommendations to actively participate in failure diagnosis and recovery. Rather than executing predefined scripts, a network of collaborating agents continuously monitors for anomalies, reasons over probable root causes, weighs remediation options, and drives recovery actions through native cloud and Kubernetes interfaces—all while preserving human oversight for high-blast-radius operations.

The results speak for themselves: faster recovery times, less operational toil, and more consistent incident response—even under pressure. More broadly, this work charts a path for how agentic AI transitions from advisory assistant to active SRE co-pilot, powering a new generation of autonomous, self-healing cloud operations.

If you're looking to move beyond reactive firefighting and toward incident workflows that are intelligent, auditable, and genuinely resilient—this session is for you.

DevopsCon New York 2026 Upcoming

AI-Powered Service Operations: Intuit's Incident Management Story

September 2026 New York City, New York, United States

KCD SF Bay Area 2026 Sessionize Event Upcoming

September 2026 San Francisco, California, United States

IEEE Cloud Summit 2026

EWOK: A Flexible Agent-Based Disaster Recovery Orchestration System for Stateful Cloud Services

June 2026 Washington, District of Columbia, United States

IEEE SVCC 2026

One Engine to Rule Them All: How We Turned a Reliability Framework into a Security Powerhouse

June 2026 San Jose, California, United States

AI Con 2026

From Reactive to Proactive: Intuit's Journey in AI-Powered Incident Management

June 2026 Seattle, Washington, United States

PagerDuty on Tour 2026

From Reactive to Proactive: AI-Powered Incident Management at Enterprise Scale with Twilio and Intuit

May 2026 San Francisco, California, United States

SRE Day 2026

From Compute to Datastores: Predictable Region Failover for Microservice Stacks

April 2026 San Francisco, California, United States

All Things AI 2026

AI-Driven Disaster Recovery for Cloud-Native Systems

March 2026 Durham, North Carolina, United States

IEEE SoutheastCon 2026

AIDR-Cloud: An Agentic AI–Driven Incident Response Framework

March 2026 Huntsville, Alabama, United States

DeveloperWeek 2026 Sessionize Event

February 2026 San Jose, California, United States

Conf42 Devops 2026

The Multi‑Workload Challenge: Predictable Failover Across the Stack

January 2026 San Jose, California, United States

API World 2025

Scalable Region Failover for Microservices: Platform-Level Resilience at Intuit

September 2025 Santa Clara, California, United States

Elastic ECK Roundtable 2022

Adobe’s journey on adopting ECK stack for building realtime observability

May 2022 San Jose, California, United States

Akshay Pratinav

Intuit, Senior Staff Software Engineer

Mountain View, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top