Session

Your Retry Is the Outage

The database recovered. Everything came back. Ninety seconds later it went down again, harder, and stayed down.

Nothing new had broken. Every client in the system had been retrying politely for two minutes, and the moment the service could accept connections again they all arrived at once. We had built a load generator by accident and pointed it at the weakest thing we owned.

Retries are the most reasonable looking code in your codebase and one of the most common causes of an incident going from bad to catastrophic. A transient failure becomes three requests instead of one. Three layers of retries stacked on top of each other becomes twenty seven. All of them fire during the exact window when the system has the least capacity to absorb them.

This talk covers what to do instead. Exponential backoff and why jitter is not optional. Timeout budgets that shrink as the call chain deepens, so the caller does not give up while three services are still working on its behalf. Circuit breakers and the failure mode where a badly tuned one causes its own outage. Backpressure, load shedding, and the counterintuitive idea that refusing work quickly is a feature.

Real incidents from a platform where an unavailable service means somebody's utility data does not land that night.

Chris Houdeshell

VP of Eng. and Ops | Bit Herder | ☕

State College, Pennsylvania, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top