Rajat Shah

Rajat Shah

Staff Software Engineer, AI Platform, Netflix

San Francisco, California, United States

Actions

Rajat is a Staff Software Engineer at Netflix, leading the technical architecture for the global ML Model Serving Infrastructure. With nearly a decade at Netflix and Amazon, he has specialized in building highly available distributed systems and stable, usable ML platforms that power personalization, search, and commerce at scale. A recipient of Amazon’s prestigious "Just Do It" Award from Jeff Bezos for his bias for action, Rajat excels at abstracting the complexities of distributed computing to drive developer velocity and platform reliability. He holds a Master's in Computer Science from North Carolina State University, and has spent his career applying that foundation to massive-scale infrastructure challenges.

Area of Expertise

  • Information & Communications Technology

Topics

  • Scalable Distributed Systems
  • Machine Learning Engineering

The Autonomous Performance Agent: A Netflix Production Story

At Netflix, performance waste is everywhere- and almost no one is looking for it.

Degradation is silent. It compounds. The manual cost of closing the loop (profile, analyze, trace, fix, validate) means most inefficiencies quietly burn compute for months before anyone acts. By the time a human gets there, the damage is done.

We decided the loop should close itself. We built an autonomous agent that continuously hunts performance inefficiencies across live production services, traces them to source code, proposes fixes, and validates results through canary deployment- grounding every decision in measured production outcomes, not model confidence.

In this talk, we'll share what it actually took to make an autonomous agent trustworthy enough to act in production: where it earns autonomy, where it doesn't, and a novel approach that changed how we think about agent reliability entirely. One finding the agent surfaced- caught, fixed, and canary-confirmed- with no ticket, no oncall, and no performance engineer in the loop.

Routing ML Inference Traffic at Internet Scale

Learn how traffic routing for ML Models poses unique challenges at internet scale, and how Netflix solves them through specialized software design patterns. Learn how various API Gateway solutions (REST), Service Mesh, and custom design can help scale traffic routing for efficient ML model serving. We will also briefly go into how various experiences on Netflix.com are powered by unique ML Models.
Here is the recent official blog post on this topic that I recently co-authored at Netflix.
https://netflixtechblog.com/state-of-routing-in-model-serving-16e22fe18741

P99 CONF 2026 Sessionize Event Upcoming

October 2026

WeAreDevelopers World Congress 2026 - North America Sessionize Event

September 2026 San Jose, California, United States

Internet Day San Francisco 2026

Routing ML Inference Traffic at Internet Scale

Learn how traffic routing for ML Models poses unique challenges at internet scale, and how Netflix solves them through specialized software design patterns. Learn how various API Gateway solutions (REST), Service Mesh, and custom design can help scale traffic routing for efficient ML model serving. We will also briefly go into how various experiences on Netflix.com are powered by unique ML Models.
Here is the recent official blog post on this topic that I recently co-authored at Netflix.

May 2026 San Francisco, California, United States

Rajat Shah

Staff Software Engineer, AI Platform, Netflix

San Francisco, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top