Session

Flight Recorder: Go's Black Box for Production

Debugging production latency is frustrating because you need to start tracing "before" the problem happens. You can't predict when a request will take 5 seconds instead of 50 milliseconds. By the time you notice, the interesting execution is gone.


Main idea: FlightRecorder continuously traces execution into a circular buffer with 1-2% overhead. When something interesting happens (slow request, error spike), you snapshot the buffer. The cause is already captured—no need to reproduce or predict when to start tracing.

Questions this talk answers:

Why can't pprof explain latency? pprof samples what's currently running on CPU. A goroutine blocked on a slow database call isn't running—it's invisible. pprof shows the CPU is idle. FlightRecorder captures blocking events, mutex contention, syscalls, and goroutine state changes—everything needed to understand where time went.

What's the overhead? 1-2% CPU overhead when recording. The runtime already collects most trace events; FlightRecorder just keeps them in a ring buffer rather than discarding them. Writes are batched and occur on dedicated goroutines to minimize latency impact.

What are MinAge and MaxBytes? MinAge says "try to keep at least this much history" (e.g., 30 seconds). MaxBytes says "but don't use more than this much memory" (e.g., 50MB). These are hints—the runtime balances them based on event rate. High-frequency events fill the buffer faster.

Why only one FlightRecorder per process? The trace infrastructure is process-global. Multiple recorders would need to coordinate buffer management and could interfere with each other. The limitation simplifies the implementation and avoids subtle bugs.

When should I snapshot? When request latency exceeds P99 threshold. When error rate spikes. When anomaly detection fires. Via a '/debug/snapshot' endpoint for manual investigation. The key insight: snapshot *after* detecting the problem, not before.

*Benefits for listeners: Attendees will know how to integrate FlightRecorder into production services, configure it appropriately (trading history length vs. memory), implement automatic snapshot triggers, and analyze traces with 'go tool trace'. The talk includes production patterns from real deployments.

This talk is currently proposed as intermediate level, but I'm open to adapting the depth if it benefits the conference.

Alex Rios

Principal Engineer @Memed

Curitiba, Brazil

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top