Category: Ecosystem and tooling

A 404 from net/http costs 62 allocations

A static route through net/http's mux costs 85.6 ns and nothing on the heap. A path that matches no pattern costs 2125.5 ns and 62 allocations.

A request flows from path matching to a highlighted matchingMethods block with method loops and tree walks, ending in a 404 response.

TL;DR

  1. Dispatch is not where anyone's latency went. On a 200-route table, gin matched a static path in 23.6 ns, echo in 26.5 ns and `net/http` in 85.6 ns — against 237.1 ns for one `json.Marshal` of a five-field response body.
  2. A path matching no pattern is the exception. `net/http` answered one in 2125.5 ns and 62 allocations, 25 times its own static match and nine times the JSON encode.
  3. The cost is `ServeMux.matchingMethods`, which walks the routing tree once per registered method — and then does it again with a trailing slash appended — so the mux can tell 404 from 405.
  4. It tracks methods, not routes. Ten times the table cost 25% more; four methods instead of one cost 56% more and 22 more allocations across the three added methods.
  5. Two identical `ServeMux` instances built from the same routes differed by up to 8.05% at p=0.000. That is the error bar every other percentage here has to clear.

A request that matched none of my 200 routes cost 2125.5 ns and 62 allocations. The same mux answered a static route in 85.6 ns and allocated nothing.

Router benchmarks are the most-published and least-trustworthy numbers in the Go ecosystem, mostly because the handler does real work and drowns the thing being measured. So this one has no handler at all — and a ruler beside the table, so the ranking can be read as a fraction of a request rather than as a league position.

What is being measured

All numbers: go1.27.0 darwin/arm64, Apple M4 Pro, 12 cores, macOS 26.6.2. Benchmarks with -benchmem -count=10, summarised by benchstat. Routers are chi v5.3.2, gin v1.12.0 and echo v4.15.4. The harness is routers/.

One route table of 200 patterns, generated from twenty resources — collection, item, nested and action shapes across four methods — and rendered into each router’s own syntax. Every route carries the same empty handler. The ResponseWriter keeps nothing: httptest.NewRecorder allocates a body buffer and a header map per request, and that would appear in every row as the recorder’s cost rather than the router’s.

chi.NewRouter() and gin.New() are built with no middleware. gin.Default() installs a logger that formats a line per request, and it would be the largest number in the table by a wide margin.

Three requests: a static path, a path carrying three parameters, and a path that matches nothing.

Dispatch, against a ruler

$ go test -run=^$ -bench='Warmup|Dispatch' -benchmem -count=10 ./routers |
    grep -v Warmup | benchstat -row /path -col /router -
        │    stdlib     │ stdlib-control │     chi      │     gin     │    echo     │
        │    sec/op     │     sec/op     │    sec/op    │   sec/op    │   sec/op    │
static      85.57n ± 2%      82.75n ± 2%   125.65n ± 1%   23.63n ± 1%   26.47n ± 1%
param      279.65n ± 1%     264.95n ± 1%   323.55n ± 3%   45.61n ± 1%   59.77n ± 4%
miss      2125.50n ± 1%    1954.50n ± 1%   298.80n ± 1%   61.60n ± 1%  435.05n ± 1%

        │  stdlib   │   chi    │   gin    │   echo   │
        │   B/op    │   B/op   │   B/op   │   B/op   │
static     0.0 ± 0%   368 ± 0%   0.0 ± 0%   0.0 ± 0%
param    112.0 ± 0%   704 ± 0%   0.0 ± 0%   0.0 ± 0%
miss    1649.0 ± 0%   416 ± 0%   0.0 ± 0%   456 ± 0%

stdlib-control is a second ServeMux built from the same 200 routes. Its true difference from stdlib is zero, so whatever benchstat prints for it is this run’s error bar — the method from benchstat or it didn’t happen. It printed −3.30%, −5.26% and −8.05%, all at p=0.000. That 8.05% is the widest control gap this site has measured, and it means no comparison below about 8% on the miss row is a result. The mutex-versus-channel sweep is where that mattered — a 1.37% win there was smaller than the control and did not survive. Here it does not bite, because the gaps in this table are factors rather than percentages.

Now the ruler. One json.Marshal of a five-field response body — an id, two strings, a timestamp and a bool — is the cheapest useful thing a handler can do after the router is finished:

Encode-12   237.1n ± 3%   256 B/op   3 allocs/op

So gin’s static dispatch is a tenth of a JSON encode, echo’s is a ninth, and net/http’s is about a third. The span from the fastest router to the slowest on a static match — gin’s 23.6 ns to chi’s 125.7 ns — is 102 ns, which is 43% of the cost of encoding the answer, before any database, any log line, or any of the work the handler exists to do. The likely finding was that dispatch is not where anyone’s latency went, and that is what the table says.

Two things are still worth reading off it. chi allocates on every dispatch, including a static one — 368 B and two allocations, which is the RouteContext it puts in the request context. And net/http allocates only when the pattern has wildcards to record: zero on the static route, 112 B on the one with three parameters.

Then there is the last row.

The row nobody benchmarks

net/http answers a path that matches no pattern in 2125.5 ns and 62 allocations. That is 25 times its own static match, and nine times the JSON encode that would have been the whole answer.

Every other router in the table treats a miss as ordinary: gin does it in 61.6 ns with nothing on the heap, chi in 298.8 ns, echo in 435.1 ns. Only the standard library turns failure into the most expensive thing it does.

I have not found a published router benchmark that measures this. The canonical suite everyone forks states its own toolchain as go1.3rc1 and describes http.ServeMux as “limited to static routes and does not support parameters” — a router Go 1.22 replaced. Gin’s own benchmarks were refreshed in March 2026 on the same class of machine as this one, and they measure matching only.

What the mux is doing

An allocation profile names it immediately:

$ go test -run=^$ -bench='Dispatch/router=stdlib/path=miss' -benchmem \
    -memprofile=/tmp/stdlib-miss.prof -memprofilerate=1 ./routers
$ go tool pprof -top -nodecount=6 -sample_index=alloc_objects /tmp/stdlib-miss.prof
File: routers.test
Type: alloc_objects
Showing nodes accounting for 1633871, 97.54% of 1675075 total
Dropped 254 nodes (cum <= 8375)
Showing top 6 nodes out of 23
      flat  flat%   sum%        cum   cum%
   1458324 87.06% 87.06%    1458324 87.06%  net/http.(*routingNode).matchPath
     54014  3.22% 90.28%      54014  3.22%  net/textproto.MIMEHeader.Set (inline)
     40514  2.42% 92.70%      40514  2.42%  slices.AppendSeq[go.shape.[]go.shape.string,go.shape.string] (inline)
     27007  1.61% 94.32%      81097  4.84%  net/http.Error
     27006  1.61% 95.93%     459102 27.41%  net/http.(*ServeMux).matchOrRedirect
     27006  1.61% 97.54%    1120754 66.91%  net/http.(*ServeMux).matchingMethods

Two thirds of the allocations are in matchingMethods, which is not the lookup. The lookup is matchOrRedirect, and it is the 27%.

The mux does not stop when nothing matches. It wants to answer 405 Method Not Allowed rather than 404 when the path exists under a different method, so it goes back and asks which methods would have matched (server.go):

// Not Found and Method Not Allowed, see if there is another pattern that
// matches except for the method.
allowedMethods := mux.matchingMethods(host, path)
if len(allowedMethods) > 0 {
	return HandlerFunc(func(w ResponseWriter, r *Request) {
		w.Header().Set("Allow", strings.Join(allowedMethods, ", "))
		Error(w, StatusText(StatusMethodNotAllowed), StatusMethodNotAllowed)
	}), "", nil, nil
}
return NotFoundHandler(), "", nil, nil

matchingMethods walks the tree, and then walks it again with a trailing slash appended, because matchOrRedirect would have tried that too. The walk itself is per method (routing_tree.go): n.children.eachPair(func(method string, c *routingNode) bool { ... c.matchPath(path, nil) ... }). One full path match per registered method, twice over.

That is the 87% sitting in matchPath.

Which axis it moves along

If the cost is a per-method walk, then it should track methods and not routes. Both sweeps are in the harness, and they agree.

Ten times the table, methods held at four:

routes │    stdlib      │     gin
   020   1522.00n ± 1%    50.35n ± 1%
   050   1642.00n ± 6%    50.01n ± 1%
   100   1684.00n ± 1%    50.48n ± 1%
   200   1904.00n ± 1%    62.18n ± 0%

The same paths, varying only how many distinct methods are registered:

methods │    stdlib     │ stdlib allocs │    gin
      1   1215.00n ± 2%       40.00 ± 0%   61.92n ± 0%
      2   1570.00n ± 3%       50.00 ± 0%   61.87n ± 0%
      3   1721.00n ± 2%       56.00 ± 0%   61.68n ± 0%
      4   1894.50n ± 1%       62.00 ± 0%   61.45n ± 0%

A ten-fold table costs 25% more. Four methods instead of one costs 56% more and 22 more allocations across three added methods, which is the eachPair loop showing up as arithmetic. gin does not move on either axis.

Note the floor: even with a single method registered, a miss is 1215 ns and 40 allocations. Most of the cost is not the sweep; it is the second pass existing at all.

What to measure instead

The ranking these benchmarks produce is real and almost never actionable. The two numbers that are:

Your dispatch against your own encoding. A router that is three times faster than another is three times faster at a tenth of what serialising the answer costs. Put one json.Marshal of your actual response body in the same table and read the ratio.

Your miss path. If net/http is your router and 404s are a meaningful share of your traffic — a scanner sweeping paths, a stale client, a health check pointed at the wrong URL — that share costs 25 times what your served requests cost to route, and it allocates where nothing else does.

What this does not measure

The handler is empty, which is the point, and it also means none of this describes a request. Middleware is absent for the same reason: chi.NewRouter() with a real stack and gin.Default() with its logger are different measurements, and in a real service they are almost certainly the larger ones.

Fiber is missing because it is built on fasthttp and does not implement http.Handler, so putting it on this bench would change what the bench is.

The request is built once and reused. Every router here either leaves it untouched or replaces the fields it needs per call, so reuse costs nothing a fresh request would not — but a server allocates a new one per connection read, and that cost is outside this table.

And the miss path measured is a path that matches nothing at all. The other miss — a path that exists under a different method, where matchingMethods finds something and the mux answers 405 — will do the same walk and then allocate an Allow header on top of it. I have not measured that one.

Frequently asked

Is net/http's ServeMux slower than gin or chi?

For matching, by a margin that almost never matters: 85.6 ns against gin's 23.6 ns on a static route, when a single json.Marshal of the response body costs 237.1 ns. For failing to match, by a margin that can: 2125.5 ns and 62 allocations against gin's 61.6 ns and zero. If your service answers a meaningful number of requests with 404, that is the row to look at.

Why does a 404 allocate at all?

Because the mux does not stop at "no handler". Before answering 404 it calls matchingMethods, which walks the routing tree once for every method registered under that node — and then walks it a second time with a trailing slash appended — collecting the set of methods that would have matched, so it can answer 405 Method Not Allowed instead when one exists. An allocation profile puts 87% of the allocations in routingNode.matchPath, reached from that pass.

Does adding routes make the miss slower?

Barely. Going from 20 routes to 200 — ten times the table — moved the miss from 1522 ns to 1904 ns, about 25%. Going from one registered method to four, with the paths unchanged, moved it from 1215 ns to 1894.5 ns and from 40 allocations to 62. The number of distinct methods is the axis; the number of routes mostly is not.

Should I switch routers because of this?

Nothing here is a reason to. The gap on a matching request is a fraction of one JSON encode, and middleware, binding and error handling — the reasons people actually choose gin or chi — are not measured at all. The useful action is smaller: measure your own dispatch next to your own response encoding, and measure a path that matches nothing.

Nolan Keir

Systems-minded Go engineer

Nolan Keir writes about Go, backend engineering, and the systems behind production software. His work focuses on concurrency, runtime behavior, performance, tooling, and the trade-offs hidden behind clean abstractions. He prefers reproducible experiments and measurable behavior over rules of thumb. He writes at Gopheria.

More about the author

Arrow keys to move, Enter to open.