Session
The Corpse Knows: Debugging Production with Process Dumps
The bug happens once a week, always around 2am, always on one node, and it never reproduces anywhere else. Your logs say nothing useful. Your APM shows a flat line and then a gap. You have added logging twice and caught nothing both times.
Stop guessing. Take a dump.
A process dump is a full copy of what your process was thinking at the moment you grabbed it. Every thread, every stack, every object on the heap, every lock and who is holding it. It is the closest thing we have to an autopsy, and it answers questions logging cannot, because you did not know to log them.
I will show you how to capture one on Windows and on Linux, including the part everyone gets stuck on, which is getting a dump out of a container before the orchestrator throws the body away. Then we open real dumps from real incidents and find real bugs. A deadlock where two threads are politely waiting for each other forever. Thread pool starvation that looked exactly like a slow database and cost us a week. A memory leak with a retention path pointing straight at an event handler nobody unsubscribed.
We will use dotnet-dump, WinDbg, and lldb. The techniques carry across languages and runtimes even if your stack is not on that list.
This is the debugging skill with the worst reputation and the best payoff. It is not hard. It is unfamiliar, which is different.
Chris Houdeshell
VP of Eng. and Ops | Bit Herder | ☕
State College, Pennsylvania, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top