Category: Concurrency in practice

The mutex was 68 times faster, until I added work

One goroutine: 2.5 ns for the mutex, 172.6 ns for the channel. Sixteen goroutines with real work: the channel wins by less than the harness's own error.

Two stacked timelines compare "sync" and "channels", with a long highlighted "channels" block taking more time before work.

TL;DR

  1. At one goroutine, incrementing a counter cost 2.5 ns behind a `sync.Mutex`, 88.9 ns through an unbuffered channel and 172.6 ns through a channel that acknowledges — factors of 35 and 68.
  2. With an empty critical section there is no crossover at any contention level from 1 to 64 goroutines. The mutex wins every row.
  3. With 146 ns of work inside the critical section the curves meet at eight goroutines and stay met. At sixteen the acknowledging channel came back 1.37% faster at p=0.000.
  4. In the same table a second copy of the mutex, compared against itself, came back 2.99% faster at p=0.000. The channel's win is smaller than the harness's own error.
  5. The mutex's cost is not monotonic in contention. It peaks at 108.9 ns with eight goroutines, then falls back and stays between 92 and 97 ns from 16 to 64.

Incrementing a counter cost 2.5 ns behind a sync.Mutex and 172.6 ns through a channel. Then I put 146 ns of work inside the critical section, and at sixteen goroutines the channel came back 1.37% faster at p=0.000. In the same run, a second copy of the mutex came back 2.99% faster than itself.

“Don’t communicate by sharing memory, share memory by communicating” is a design proverb — Rob Pike’s, from the Go Proverbs talk at Gopherfest SV 2015 — and its second half gets quoted as a performance claim. This post draws the curve the claim implies — one counter, five ways of guarding it, from one goroutine to sixty-four — and then checks whether the place the curves cross is a contention level or a rounding error.

What is being measured

All numbers: go1.27.0 darwin/arm64, Apple M4 Pro, 12 cores, macOS 26.6.2. Benchmarks with -benchmem -count=10, summarised by benchstat. The harness is mutexchan/.

Five implementations of one counter, behind an interface that hands each goroutine its own closure — so per-goroutine state, such as a reply channel, is paid for once per goroutine and not once per operation:

type Counter interface {
	Handle() func()
	Close() int64
}

mutex takes a sync.Mutex. atomic is a single atomic.Int64 add, which has no critical section at all and is here as a floor. chan-unbuf and chan-buf1024 send to an owner goroutine and do not wait. chan-ack sends and waits for the acknowledgement, which is the only channel shape that offers what a Lock/Unlock pair offers: the increment has happened when the call returns.

Close drains the channel and compares the total against b.N, so an asynchronous implementation cannot look fast by leaving work undone.

The curve with an empty critical section

$ go test -run=^$ -bench='Warmup|Counter' -benchmem -count=10 ./mutexchan |
    grep -v Warmup | benchstat -row /goroutines -col /impl -
        │    mutex     │  mutex-control  │    atomic    │  chan-unbuf   │ chan-buf1024  │   chan-ack    │
        │    sec/op    │     sec/op      │    sec/op    │    sec/op     │    sec/op     │    sec/op     │
1          2.549n ± 1%      2.540n ±  2%   1.828n ±  0%    88.910n ± 1%   24.480n ±  3%  172.600n ± 5%
2         19.375n ± 5%     20.090n ± 14%   6.269n ±  9%   124.300n ± 2%   28.985n ±  2%  175.400n ± 4%
4         68.730n ± 1%     73.640n ± 14%  14.750n ±  3%   201.450n ± 5%   38.000n ± 11%  171.450n ± 2%
8        108.850n ± 2%    108.000n ±  0%  23.720n ± 22%   271.300n ± 5%   51.370n ±  1%  173.400n ± 2%
16        94.400n ± 0%     93.880n ±  3%  43.810n ±  4%   305.550n ± 3%   69.460n ±  9%  176.450n ± 1%
32        96.590n ± 0%     96.360n ±  1%  43.330n ±  5%   404.300n ± 3%  110.450n ±  1%  175.250n ± 2%
64        91.510n ± 1%     91.270n ±  2%  41.350n ± 36%   502.400n ± 4%  188.300n ±  5%  183.700n ± 2%
geomean   43.780n          44.310n        16.370n         233.200n        57.190n        175.400n

Every row reported 0 B/op and 0 allocs/op, so the rest of this post is about time.

The second column is not a fifth implementation. mutex-control is another mutex — same type, same code, registered as its own row and measured in the same run — so its true difference from mutex is zero by construction, and whatever benchstat prints for it is this session’s error bar. That is the method from benchstat or it didn’t happen, and it is the only reason any percentage below means anything.

There is no crossover. The mutex wins every level from one goroutine to sixty-four; the unbuffered channel starts 35 times behind and ends 5.5 times behind, the acknowledging one starts 68 times behind and ends twice behind. The gap narrows and never closes.

The curve with 146 ns in it

An empty critical section is not the case anybody has. BenchmarkSpin prices 128 xorshift iterations at 145.9 ns; that goes inside the lock, and inside the owner goroutine’s loop, so both designs serialise the same work:

$ go test -run=^$ -bench='Warmup|Counter' -benchmem -count=10 ./mutexchan \
    -spin=128 | grep -v Warmup | benchstat -row /goroutines -col /impl -
        │    mutex    │ mutex-control │           chan-unbuf            │          chan-buf1024           │            chan-ack            │
        │   sec/op    │    sec/op     │   sec/op     vs base            │    sec/op     vs base           │   sec/op     vs base           │
1         147.4n ± 0%    147.5n ± 0%   358.2n ± 3%  +143.05% (p=0.000)   209.6n ±  5%  +42.20% (p=0.000)   404.1n ± 8%  +174.15% (p=0.000)
2         230.7n ± 1%    236.3n ± 2%   304.5n ± 2%   +32.02% (p=0.000)   305.4n ±  8%  +32.43% (p=0.000)   357.0n ± 4%   +54.78% (p=0.000)
4         266.1n ± 1%    265.3n ± 2%   445.2n ± 7%   +67.27% (p=0.000)   340.6n ±  4%  +27.95% (p=0.000)   339.1n ± 1%   +27.39% (p=0.000)
8         322.6n ± 1%    321.9n ± 1%   415.2n ± 2%   +28.72% (p=0.000)   359.6n ± 10%  +11.47% (p=0.000)   327.1n ± 1%    +1.40% (p=0.000)
16        332.1n ± 1%    331.7n ± 2%   471.9n ± 2%   +42.06% (p=0.000)   419.1n ±  4%  +26.18% (p=0.000)   327.6n ± 1%    -1.37% (p=0.000)
32        332.6n ± 0%    322.7n ± 2%   516.1n ± 7%   +55.16% (p=0.000)   497.9n ± 15%  +49.69% (p=0.000)   328.2n ± 1%    -1.32% (p=0.001)
64        331.8n ± 0%    318.9n ± 4%   508.6n ± 3%   +53.32% (p=0.000)   566.2n ±  9%  +70.66% (p=0.000)   333.9n ± 8%         ~ (p=0.160)

Here is the crossover the proverb implies. At one goroutine the acknowledging channel is still 174% behind. By eight it is 1.40% behind, by sixteen it is 1.37% ahead at p=0.000, and it stays ahead at thirty-two.

A 1.37% win at p=0.000 is the kind of number a pull request gets approved on. So I read the column next to it.

The copy of the mutex that beat the mutex

mutex-control runs the same code as mutex. Its true delta is zero. At thirty-two goroutines benchstat reported it 2.99% faster at p=0.000, and at sixty-four 3.87% faster at p=0.001.

Both of those are larger than every win the channel produced. The channel’s 1.37% at sixteen goroutines is less than half of what a literal copy of the baseline managed against itself in the same run, at a comparable p-value. The p-value is not lying: those two sets of samples really did come from different distributions. They came from different distributions because they ran in different positions, on different cores, at different points in the process’s life — and none of that is the code.

The honest reading of the busy table is therefore not “the channel wins past sixteen goroutines”. It is that past eight goroutines this benchmark can no longer tell the two apart, and it says so with a number: about 3%. That number is specific to this machine and this session, which is exactly why it is measured rather than assumed — on a quieter machine a real 1.37% might survive.

Why the channel’s line is flat

The acknowledging channel cost 172.6 ns at one goroutine and 183.7 ns at sixty-four — a 6% rise across a sixty-four-fold change in contention. Nothing else in the table behaves like that.

One goroutine owns the counter and applies every increment, so throughput is bounded by the handoff rather than by the senders. A second sender does not make the receive faster or slower; it only waits longer for its turn, and the benchmark divides that waiting back out again.

The handoff is a park and an unpark: the sender blocks, the runtime takes its goroutine off the P and runs something else, and the owner’s receive makes it runnable again — the same mechanism that lets a blocked goroutine cost nothing when the runtime owns the thing it is blocked on.

The buffered channel is the exception that confirms it. chan-buf1024 is cheap while the buffer absorbs the senders — 24.5 ns at one goroutine — and degrades steadily as they outrun the owner, reaching 188.3 ns at sixty-four. That is within 3% of what the acknowledging channel cost at every level, which is the handoff price arriving late rather than being avoided. A buffer does not remove the handoff; it defers it until the buffer is full.

Why the mutex gets faster after eight goroutines

The mutex column is not monotonic, and this surprised me. It climbs from 2.5 ns to 108.9 ns at eight goroutines, then falls to 94.4 ns at sixteen and stays between 92 and 97 ns through sixty-four.

sync.Mutex has two modes. In normal mode an arriving goroutine spins briefly and may take the lock ahead of goroutines already queued — fast when contention is low, unfair when it is high. When a waiter has been queued for more than a millisecond the mutex switches to starvation mode and hands the lock directly to the head of the queue, which removes both the spinning and the barging. The threshold is the constant starvationThresholdNs = 1e6 in src/internal/sync/mutex.go, where sync.Mutex has forwarded its implementation since Go 1.24.

Whether that is what happened here is a separate question, and the sweep cannot answer it: the mode is not observable from outside the runtime. So fairness/ measures its two symptoms — how evenly the lock was distributed, and how long the longest waiter waited.

$ go build -o /tmp/fairness ./mutexchan/fairness
$ /tmp/fairness -dur=3s
n          min ops    median ops       max ops   max/min
2         51271680      60759040      60759040      1.19
4         10675200      10770432      10778624      1.01
8          3389440       3465216       3486720      1.03
16         1918976       2013184       2026496      1.06
32          890880        928768        949248      1.07
64          487424        512000        540672      1.11

$ /tmp/fairness -dur=3s -timed
n         max wait      p99 wait        median   reached 1ms
2         15.939ms            0s          83ns   true
4          1.878ms           6µs         250ns   true
8          1.568ms          30µs         250ns   true
16           1.7ms          63µs         125ns   true
32         2.937ms         149µs         166ns   true
64         3.546ms         338µs         125ns   true

Both symptoms are there. The longest wait crosses one millisecond at every level, so the threshold is reached. And the distribution is close to flat — above two goroutines, no goroutine got more than 11% more turns than the slowest — which is what a FIFO handoff looks like and not what barging looks like. The timed run’s two time.Now calls per Lock inflate the waits, which makes the crossing a weaker claim rather than a stronger one; the distribution run has no timing in it at all.

The tail also shows this is not a permanent state. The median wait is under 250 ns and the p99 is microseconds; only the extreme tail reaches a millisecond, and the mutex returns to normal mode as soon as a waiter is served quickly. The plateau at 92–97 ns is a lock flipping between two modes, not sitting in one.

Eight goroutines on twelve cores is where the spinning stops paying and the flipping has not yet settled. It is the worst of both, and it is the peak of the curve.

Where the boundary actually is

“At what contention level does a channel beat a mutex” has no answer on this machine, and not because the answer is very large. The boundary is not a goroutine count at all. It is a comparison between two numbers you have to measure yourself: the size of the win, against the size of the win a copy of your baseline reports. If the first is smaller, the benchmark has not distinguished the implementations, whatever the p-value says — and adding the copy costs one more entry in the table and no extra command.

What this does not measure

A counter is the smallest possible critical section, and the two work levels here — nothing, and 146 ns — are both far below what a real one does. A section holding a map write or an I/O call moves both curves, and may move them differently.

There are also no other goroutines in it. A service under load has hundreds of runnable goroutines competing for the same Ps, and a design with a single owner is a different proposition when that owner has to wait its turn to run. It says nothing about sync.RWMutex, sharding or per-P accumulation either, all of which remove the contention instead of serialising it — which is the actual answer when the contention is the problem.

And fairness/ measures the symptoms of starvation mode, not the mode. An even distribution and a millisecond wait are what that mode produces, but not the only things that could produce them. Reading the flag needs the runtime’s own instrumentation, and that is a different experiment.

Frequently asked

Is a channel slower than a mutex in Go?

For guarding a shared value, yes, and by a wide margin when there is nothing inside the critical section: 2.5 ns against 172.6 ns at one goroutine on the machine below. The margin narrows as contention rises and as the critical section does real work, and at eight goroutines with 146 ns of work in it the two are no longer separable. Neither result says anything about the cases channels exist for — moving work between goroutines, cancellation, fan-out — which this benchmark does not measure.

At what contention level do channels become faster than mutexes?

On this machine, none that can be defended. With an empty critical section the mutex won every row from 1 to 64 goroutines. With work inside it, the acknowledging channel came back 1.37% faster at 16 goroutines with p=0.000 — but a second copy of the mutex, measured against itself in the same run, came back 2.99% faster with p=0.000. A win smaller than the control pair is a property of the harness, not of the code.

Why does the channel-owned counter cost the same at 1 goroutine as at 64?

Because one owner goroutine applies every increment, so throughput is bounded by the handoff rather than by the senders. Adding senders does not make the handoff slower; it makes each sender wait longer, and the benchmark divides that wait back out again. The measured range across a 64-fold change in contention was 172.6 ns to 183.7 ns.

Should I replace a mutex with a channel for performance?

Not on the strength of a number smaller than your own control pair. Put a second copy of the baseline in the same benchmark table, read the win against what that copy reports, and if the win is smaller, the measurement cannot tell the two apart. The reasons to choose a channel are ownership and composition, and those are not settled by nanoseconds.

Nolan Keir

Systems-minded Go engineer

Nolan Keir writes about Go, backend engineering, and the systems behind production software. His work focuses on concurrency, runtime behavior, performance, tooling, and the trade-offs hidden behind clean abstractions. He prefers reproducible experiments and measurable behavior over rules of thumb. He writes at Gopheria.

More about the author

Arrow keys to move, Enter to open.