# A 404 from net/http costs 62 allocations

> A static route through net/http's mux costs 85.6 ns and nothing on the heap. A path that matches no pattern costs 2125.5 ns and 62 allocations.

- Published: 2026-09-27
- Tags: net-http, benchmarking, performance, stdlib
- Source: https://gopheria.com/blog/http-router-dispatch-cost/
- Language: en-US
- Author: Nolan Keir

---
A request that matched none of my 200 routes cost 2125.5 ns and 62 allocations.
The same mux answered a static route in 85.6 ns and allocated nothing.

Router benchmarks are the most-published and least-trustworthy numbers in the Go
ecosystem, mostly because the handler does real work and drowns the thing being
measured. So this one has no handler at all — and a ruler beside the table, so
the ranking can be read as a fraction of a request rather than as a league
position.

## What is being measured

All numbers: `go1.27.0 darwin/arm64`, Apple M4 Pro, 12 cores, macOS 26.6.2.
Benchmarks with `-benchmem -count=10`, summarised by `benchstat`. Routers are
`chi v5.3.2`, `gin v1.12.0` and `echo v4.15.4`. The harness is
[`routers/`](https://github.com/CognatePress/experiments.gopheria.com/tree/main/routers).

One route table of 200 patterns, generated from twenty resources — collection,
item, nested and action shapes across four methods — and rendered into each
router's own syntax. Every route carries the same empty handler. The
`ResponseWriter` keeps nothing: `httptest.NewRecorder` allocates a body buffer
and a header map per request, and that would appear in every row as the
recorder's cost rather than the router's.

`chi.NewRouter()` and `gin.New()` are built with no middleware. `gin.Default()`
installs a logger that formats a line per request, and it would be the largest
number in the table by a wide margin.

Three requests: a static path, a path carrying three parameters, and a path that
matches nothing.

## Dispatch, against a ruler

```text
$ go test -run=^$ -bench='Warmup|Dispatch' -benchmem -count=10 ./routers |
    grep -v Warmup | benchstat -row /path -col /router -
```

```text
        │    stdlib     │ stdlib-control │     chi      │     gin     │    echo     │
        │    sec/op     │     sec/op     │    sec/op    │   sec/op    │   sec/op    │
static      85.57n ± 2%      82.75n ± 2%   125.65n ± 1%   23.63n ± 1%   26.47n ± 1%
param      279.65n ± 1%     264.95n ± 1%   323.55n ± 3%   45.61n ± 1%   59.77n ± 4%
miss      2125.50n ± 1%    1954.50n ± 1%   298.80n ± 1%   61.60n ± 1%  435.05n ± 1%

        │  stdlib   │   chi    │   gin    │   echo   │
        │   B/op    │   B/op   │   B/op   │   B/op   │
static     0.0 ± 0%   368 ± 0%   0.0 ± 0%   0.0 ± 0%
param    112.0 ± 0%   704 ± 0%   0.0 ± 0%   0.0 ± 0%
miss    1649.0 ± 0%   416 ± 0%   0.0 ± 0%   456 ± 0%
```

`stdlib-control` is a second `ServeMux` built from the same 200 routes. Its true
difference from `stdlib` is zero, so whatever `benchstat` prints for it is this
run's error bar — the method from
[benchstat or it didn't happen](/blog/benchstat-or-it-didnt-happen/). It printed
**−3.30%, −5.26% and −8.05%, all at `p=0.000`**. That 8.05% is the widest control
gap this site has measured, and it means no comparison below about 8% on the
miss row is a result. The
[mutex-versus-channel sweep](/blog/mutex-versus-channel-at-contention/) is where
that mattered — a 1.37% win there was smaller than the control and did not
survive. Here it does not bite, because the gaps in this table are factors
rather than percentages.

Now the ruler. One `json.Marshal` of a five-field response body — an id, two
strings, a timestamp and a bool — is the cheapest useful thing a handler can do
after the router is finished:

```text
Encode-12   237.1n ± 3%   256 B/op   3 allocs/op
```

So gin's static dispatch is a tenth of a JSON encode, echo's is a ninth, and
`net/http`'s is about a third. The span from the fastest router to the slowest
on a static match — gin's 23.6 ns to chi's 125.7 ns — is 102 ns, which is 43% of
the cost of encoding the answer, before any database, any log line, or any of
the work the handler exists to do. The likely finding was that dispatch is not
where anyone's latency went, and that is what the table says.

Two things are still worth reading off it. chi allocates on **every** dispatch,
including a static one — 368 B and two allocations, which is the `RouteContext`
it puts in the request context. And `net/http` allocates only when the pattern
has wildcards to record: zero on the static route, 112 B on the one with three
parameters.

Then there is the last row.

## The row nobody benchmarks

`net/http` answers a path that matches no pattern in 2125.5 ns and 62
allocations. That is **25 times its own static match**, and nine times the JSON
encode that would have been the whole answer.

Every other router in the table treats a miss as ordinary: gin does it in 61.6 ns
with nothing on the heap, chi in 298.8 ns, echo in 435.1 ns. Only the standard
library turns failure into the most expensive thing it does.

I have not found a published router benchmark that measures this. The canonical
suite everyone forks states its own toolchain as `go1.3rc1` and describes
`http.ServeMux` as "limited to static routes and does not support parameters" —
a router Go 1.22 replaced. Gin's own benchmarks were refreshed in March 2026 on
the same class of machine as this one, and they measure matching only.

## What the mux is doing

An allocation profile names it immediately:

```text
$ go test -run=^$ -bench='Dispatch/router=stdlib/path=miss' -benchmem \
    -memprofile=/tmp/stdlib-miss.prof -memprofilerate=1 ./routers
$ go tool pprof -top -nodecount=6 -sample_index=alloc_objects /tmp/stdlib-miss.prof
File: routers.test
Type: alloc_objects
Showing nodes accounting for 1633871, 97.54% of 1675075 total
Dropped 254 nodes (cum <= 8375)
Showing top 6 nodes out of 23
      flat  flat%   sum%        cum   cum%
   1458324 87.06% 87.06%    1458324 87.06%  net/http.(*routingNode).matchPath
     54014  3.22% 90.28%      54014  3.22%  net/textproto.MIMEHeader.Set (inline)
     40514  2.42% 92.70%      40514  2.42%  slices.AppendSeq[go.shape.[]go.shape.string,go.shape.string] (inline)
     27007  1.61% 94.32%      81097  4.84%  net/http.Error
     27006  1.61% 95.93%     459102 27.41%  net/http.(*ServeMux).matchOrRedirect
     27006  1.61% 97.54%    1120754 66.91%  net/http.(*ServeMux).matchingMethods
```

Two thirds of the allocations are in `matchingMethods`, which is not the lookup.
The lookup is `matchOrRedirect`, and it is the 27%.

The mux does not stop when nothing matches. It wants to answer `405 Method Not
Allowed` rather than `404` when the path exists under a different method, so it
goes back and asks which methods *would* have matched
([`server.go`](https://github.com/golang/go/blob/go1.27.0/src/net/http/server.go#L2753-L2762)):

```go
// Not Found and Method Not Allowed, see if there is another pattern that
// matches except for the method.
allowedMethods := mux.matchingMethods(host, path)
if len(allowedMethods) > 0 {
	return HandlerFunc(func(w ResponseWriter, r *Request) {
		w.Header().Set("Allow", strings.Join(allowedMethods, ", "))
		Error(w, StatusText(StatusMethodNotAllowed), StatusMethodNotAllowed)
	}), "", nil, nil
}
return NotFoundHandler(), "", nil, nil
```

`matchingMethods` walks the tree, and then walks it again with a trailing slash
appended, because `matchOrRedirect` would have tried that too. The walk itself
is per method
([`routing_tree.go`](https://github.com/golang/go/blob/go1.27.0/src/net/http/routing_tree.go#L229-L242)):
`n.children.eachPair(func(method string, c *routingNode) bool { ... c.matchPath(path, nil) ... })`.
One full path match per registered method, twice over.

That is the 87% sitting in `matchPath`.

## Which axis it moves along

If the cost is a per-method walk, then it should track methods and not routes.
Both sweeps are in the harness, and they agree.

Ten times the table, methods held at four:

```text
routes │    stdlib      │     gin
   020   1522.00n ± 1%    50.35n ± 1%
   050   1642.00n ± 6%    50.01n ± 1%
   100   1684.00n ± 1%    50.48n ± 1%
   200   1904.00n ± 1%    62.18n ± 0%
```

The same paths, varying only how many distinct methods are registered:

```text
methods │    stdlib     │ stdlib allocs │    gin
      1   1215.00n ± 2%       40.00 ± 0%   61.92n ± 0%
      2   1570.00n ± 3%       50.00 ± 0%   61.87n ± 0%
      3   1721.00n ± 2%       56.00 ± 0%   61.68n ± 0%
      4   1894.50n ± 1%       62.00 ± 0%   61.45n ± 0%
```

A ten-fold table costs 25% more. Four methods instead of one costs 56% more and
**22 more allocations across three added methods**, which is the `eachPair` loop
showing up as arithmetic. gin does not move on either axis.

Note the floor: even with a single method registered, a miss is 1215 ns and 40
allocations. Most of the cost is not the sweep; it is the second pass existing
at all.

## What to measure instead

The ranking these benchmarks produce is real and almost never actionable. The
two numbers that are:

**Your dispatch against your own encoding.** A router that is three times faster
than another is three times faster at a tenth of what serialising the answer
costs. Put one `json.Marshal` of your actual response body in the same table and
read the ratio.

**Your miss path.** If `net/http` is your router and 404s are a meaningful share
of your traffic — a scanner sweeping paths, a stale client, a health check
pointed at the wrong URL — that share costs 25 times what your served requests
cost to route, and it allocates where nothing else does.

:::note
The 405 pass is not a bug. Answering `405` with a correct `Allow` header is
behaviour the other routers here mostly do not offer, and it has to be paid for
somewhere. It is worth knowing that the bill arrives on the failure path.
:::

## What this does not measure

The handler is empty, which is the point, and it also means none of this
describes a request. Middleware is absent for the same reason: `chi.NewRouter()`
with a real stack and `gin.Default()` with its logger are different measurements,
and in a real service they are almost certainly the larger ones.

Fiber is missing because it is built on `fasthttp` and does not implement
`http.Handler`, so putting it on this bench would change what the bench is.

The request is built once and reused. Every router here either leaves it
untouched or replaces the fields it needs per call, so reuse costs nothing a
fresh request would not — but a server allocates a new one per connection read,
and that cost is outside this table.

And the miss path measured is a path that matches nothing at all. The other
miss — a path that exists under a different method, where `matchingMethods`
finds something and the mux answers 405 — will do the same walk and then
allocate an `Allow` header on top of it. I have not measured that one.
