Session

Why Go Hides Its Spinlocks

A spinlock is the difference between pacing by the door versus taking a nap while waiting for a package. One burns CPU cycles. The other yields to the scheduler. Both have their place, but Go deliberately hides spinlocks from you.

This talk explains the core trade-off between spinning and parking, reveals the hidden spin within a Mutex, and explores why Go's runtime uses a "spinbit" design that allows only one goroutine to spin at a time.

You'll learn when spinning wins, when it's catastrophic, and why adaptive hybrid locks beat both pure approaches.


Main idea: A waiting goroutine can either spin (loop checking the lock, burning CPU) or park (yield to the scheduler, pay context switch overhead). Go's 'sync.Mutex' already spins internally (up to 4 iterations) before parking. The runtime uses the "spinbit" design: only ONE goroutine can spin at a time, preventing thundering herd problems.

Questions this talk answers:

When does spinning beat parking? Spinning wins when the critical section is shorter than a context switch (~1-2 microseconds). For sub-100ns critical sections on multi-core systems with low contention, spinning avoids scheduler overhead entirely. For anything longer, parking wins because you're just wasting cycles.

What does x86 PAUSE do? PAUSE is a hint that tells the CPU "I'm in a spin loop." It reduces power consumption, avoids memory order violations on speculative execution, and prevents the spinning core from flooding the memory bus with lock reads. Without PAUSE, spin loops can actually slow down the lock holder.

How does sync.Mutex spin? The runtime's 'runtime_canSpin' function checks: Are we on a multi-core machine? Have we spun fewer than 4 times? Is there an idle P available? Is the current goroutine's run queue empty? If all conditions are met, the goroutine spins briefly before parking.

What's the spinbit? A single bit in the mutex state that says "someone is already spinning." When set, other waiters go directly to sleep. This prevents N goroutines from all spinning simultaneously—only one burns CPU while others wait efficiently.

Why no SpinMutex in the standard library? Pure spinlocks are almost never the right choice in Go. Goroutines are cooperatively scheduled; a spinning goroutine can prevent the lock holder from running on the same P. The adaptive approach in 'sync.Mutex' handles the common cases correctly.

Attendees will understand when Mutex spins vs. parks, why pure spinlocks usually hurt Go programs, and the rare scenarios (sub-100ns critical sections, known multi-core, low contention) where custom spinning might help.

A possible beginner-level adaptation:

For a beginner audience, the talk would focus on the intuition behind spinning vs. parking rather than hardware details. I'd use the analogy throughout: spinning is pacing by the door waiting for a package, parking is taking a nap and asking to be woken up. The key insight is that both have costs: spinning wastes energy, and napping takes time to wake up. The talk would cover why sync.Mutex is almost always the right choice (it already handles the trade-off internally), demonstrate what happens when you write a naive spin loop (CPU spikes, other goroutines starve), and show the one-liner check to see if lock contention is your problem (go tool pprof mutex profile). Attendees would leave understanding why they shouldn't write their own spinlock and how to diagnose if sync.Mutex is a bottleneck in their application.

A possible advanced-level adaptation:

For an advanced audience, the talk would dive into the hardware and runtime implementation. We'd examine x86 PAUSE semantics (pipeline flush, memory-order implications, power reduction), how cache-coherence protocols (MESI) make naive spinning flood the memory bus, and the exact assembly Go generates for sync.Mutex.Lock. The talk would walk through runtime.lock2 and runtime_canSpin source code, explaining each heuristic: why 4 spin iterations, why check for idle Ps, why the run queue must be empty. We'd benchmark different spin strategies (exponential backoff, TTAS vs. TAS, ticket locks) and show when each wins. The spinbit design would get detailed treatment: how one bit prevents thundering herd, why it's set/cleared atomically with the lock state. Attendees would leave able to read and modify src/runtime/lock_futex.go and make informed decisions about custom synchronization primitives for extreme performance scenarios.

Alex Rios

Principal Engineer @Memed

Curitiba, Brazil

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top