How binder, our zero-dependency HTTP request binder for Go, cut allocations per request by up to 93% and memory by up to 90% — and the habits that keep it there.
binder maps an HTTP request onto a Go struct: path values, the query string, a JSON, form or multipart body, cookies and headers, all in one call, driven by struct tags.
type CreateOrder struct {
TeamID int `path:"team"`
DryRun bool `query:"dry_run"`
Email string `body:"email,required"`
Trace string `header:"X-Request-ID"`
Session string `cookie:"session"`
}
var req CreateOrder
if err := binder.Bind(r, &req); err != nil { ... }
A binder runs on every request a service handles, so whatever it allocates, every request pays for. That cost does not show up as latency in a microbenchmark. It shows up as GC pressure under load: more garbage per request means the collector runs more often, and the cost lands on whichever requests happen to be in flight when it does. Cutting allocations is buying tail latency and headroom, not the nanoseconds in the tables below.
Which is what made binder's starting position awkward. It was already the fastest of the reflective binders, and it allocated the most of them — 27 allocations on a request using every source, against Echo's 15 and Gin's 21. The per-type tag cache and a JSON walk that wrote straight into fields bought more time than the extra allocations cost, so the benchmark that mattered looked good while the one that predicts behaviour under load looked bad. For 1.2.0 we went through binder looking for allocations we could avoid. Here is what we changed, what it bought, and the habits that stop it creeping back.
The Result
The comparison figures below come from one interleaved run of the benchmarks in binder's repository: five fields from a query string, five from a JSON body, and a request using every source at once. The per-source figures quoted later — the type cache, form bodies, multipart — come from binder's own benchmark suite rather than this comparison. v1.1.0 and v1.2.0 were measured against the same benchmark code, alternating old, new, old, new, on an Apple M2 with Go 1.27, -benchtime=200ms -count=10, taking the median of twenty samples. Reproduce with make bench.
| Request | v1.1.0 | v1.2.0 |
|---|---|---|
| Query string, 5 fields | 690 ns, 672 B, 14 allocs | 527 ns, 64 B, 1 alloc |
| JSON body, 5 fields | 807 ns, 1,240 B, 23 allocs | 633 ns, 256 B, 8 allocs |
| Path, query, body, header and cookie | 1,134 ns, 1,928 B, 27 allocs | 614 ns, 248 B, 6 allocs |
Allocations are down 65 to 93%, memory 79 to 90%, and time 22 to 46%. Against the framework binders on the same requests, binder now makes the fewest allocations or ties. On a request using every source it takes 6 allocations to Echo's 15 and Gin's 21.
The rest of this post is how.
Measure Before You Guess
Every change below started with a benchmark and ended with one:
go test -run '^$' -bench . -benchmem -benchtime=200ms -count=10
-benchmem gives you bytes and allocations per operation. Ten runs give you a median rather than a lucky number. When you compare two versions, run them alternately, old then new then old, not one after the other. The first measurement we took after one change showed every benchmark 60% slower — including the query benchmark, which that change could not touch. Running it twice more gave the usual figures. A laptop's thermal state drifts over a run, and alternating puts that drift into both sides equally. When a benchmark you didn't touch moves, suspect the machine before the code.
To find out which line allocates, record every allocation and rank by count rather than bytes:
go test -run '^$' -bench '^BenchmarkJSON_Binder$' -benchmem \
-benchtime 200000x -memprofile mem.prof -memprofilerate 1 -o bench.test
go tool pprof -sample_index=alloc_objects -top bench.test mem.prof
-memprofilerate 1 records every allocation rather than a sample, and alloc_objects ranks by count, which is what allocs/op counts. Divide each count by the iteration count to get allocs/op.
The profile will not add up, and that is expected. Ours accounted for 16 of the 23 allocations on the JSON body. The runtime packs small pointer-free allocations — under 16 bytes, which includes short strings — into shared blocks. -benchmem counts each one; the profiler sees only the block. Short JSON member names like "age" vanish into that gap, so the profile showed two allocations where there were really eight strings. Where the numbers don't reconcile, reason from the code, or isolate the suspect with testing.AllocsPerRun.
Don't Put Strings Into Interfaces
The biggest single cost was hiding in plain sight, and it was neither reflection nor JSON. It was a function signature. binder converted values through a general extractFieldValue(...) (any, bool, error). Passing a string as any stores it in an interface, and an interface holding a string needs the string's header on the heap. That's one allocation per path, query, header or cookie field, on every request, before any conversion has happened.
So values from those four sources, which are always strings, now take a typed path that never touches any:
// Simplified. A field's kind is resolved once per type, so
// the string goes straight into it.
switch fi.Fast {
case fastString:
field.SetString(s)
case fastInt:
n, err := strconv.ParseInt(s, 10, 64)
...
}
The general path still exists for pointers, named types, slices, maps and anything implementing encoding.TextUnmarshaler; the common case just doesn't go near it. The query case lost exactly the 5 allocations predicted and got 28% faster — and only half that saving was the allocations. The rest was skipping the type switches in the general setter, which a known string never needed.
Read Only What You Asked For
r.URL.Query() parses the whole query string into a url.Values: a map, a slice per key and a string per value. That's 7 allocations to read 5 of them.
binder scans RawQuery once per field instead, stopping at the first match. A value without escapes is returned as a substring of the URL, so it costs nothing. Past 64 parameters, scanning once per field would cost more than parsing once, so binder falls back to url.ParseQuery and caches the result — which also leaves net/url's own parameter limit, added for a CVE, applied by the standard library itself.
Cookies work the same way: rather than r.Cookie(name), which builds a *Cookie for every cookie in the header so binder can read one .Value, binder scans for the one it wants. Headers are looked up by indexing r.Header with a key canonicalised once per type, because Header.Get canonicalises its argument on every call, and that allocates when the tag isn't already in canonical form.
This is the one change that cost time to save allocations. A per-field scan repeats work that one parse did once, so on the five-parameter query it measured about 4% slower — 501 ns to 522 ns, A/B'd — while removing 7 allocations. We kept it: the GC pressure is worth 20 ns here, and the mixed request, which is binder's reason to exist, got 18% faster from the same change. An "obvious" follow-up optimisation went the other way. Checking the key first and skipping the escape handling for non-matching pairs looked like a clear win, and measured 525 ns against 647 ns — worse, because the key check cost more per call than the unescape it was avoiding. It was reverted, with a comment saying why.
Reimplementing parsing from the standard library is a good way to introduce bugs, so each scanner has a differential fuzz test. It feeds random input to both the scanner and url.ParseQuery (or http.Request.Cookie) and fails if their answers ever differ. They have to agree byte for byte, including on malformed input: 3.4 million executions on the query scanner and 2.66 million on the cookie scanner, with no difference found. Both run in CI.
Stream JSON Instead of Building Maps
The old JSON path decoded the body into map[string]any, then copied values into fields. That's an allocation per member for the interface value, plus the map, plus the strings, all thrown away a moment later.
binder now walks the body's tokens with encoding/json/jsontext and writes each value straight into its field. That package is the token-level half of json/v2, in the standard library as of Go 1.27, which binder requires; there is no longer a fallback decoder for toolchains built with GOEXPERIMENT=nojsonv2.
- An unknown member is skipped without being decoded.
SkipValuechecks that it's well-formed and moves on. - A matching token sets a predeclared field directly. A JSON string into a
string, or a number into anint, never becomes anany. - Member names are looked up without allocating. The name is read as raw bytes, and the map lookup converts it in place:
The compiler recognisesi, ok := of.index[string(name)] // name is []bytem[string(b)]and doesn't allocate for the conversion. Converting first,key := string(name), would allocate for every member. Only a name written with escapes, such as"age", is decoded at all. - The decoder is pooled. A
jsontext.Decodercosts several allocations to create, so binder keeps them in async.Pool, each with the buffer it reads from. jsontext parses abytes.Bufferin place rather than copying it, which saves the read buffer and means a pooled decoder holds no copy of an old body. The buffer is cleared before the decoder goes back, so the pool never keeps a request's data alive.
Parsing in place is worth one more test than it looks. The bytes the decoder reads are the same bytes that back the restored r.Body, so if decoding ever rewrote them — while unescaping, say — a later handler would read a corrupted body. A test binds a body full of escapes three times over on reused decoders and checks the restored bytes one by one.
Pay for Reflection Once Per Type
Reflection is where binding gets expensive, but most of it describes the type, not the request. Which fields have tags, what kind each one is, the canonical form of each header name, whether a type implements encoding.TextUnmarshaler: all of that is worked out the first time a type is bound and cached. The same goes for a per-type decode plan for nested JSON, because reflect.Type.Implements is far too slow to ask about every element of an array.
A benchmark that clears the cache before every iteration shows what this is worth: 2,384 ns and 20 allocations cold, against 747 ns and 7 warm.
Keep Bookkeeping on the Stack
Binding needs scratch space: which fields were sent, which failed. It's easy to make a slice for it, which means one allocation per request — or, worse, per nested object.
- Which fields bound fits in a
[64]boolarray on the stack, for any struct of up to 64 fields, which is nearly every request type. Only wider structs get a slice. - Which members a nested object sent is a
uint64bit set. A wide struct only allocates when a member past the 64th is actually sent. - The failure list stays
niluntil something fails. A request that binds cleanly, which is most of them, never allocates for errors it doesn't have.
The per-object version of this was the worst of the three, because it scaled with the body rather than the request. Every nested object used to allocate bookkeeping sized to its struct whether anything failed or not, so 100,000 empty objects decoded into a 32-field struct cost 290 MB. A clean object now allocates nothing extra, and a regression test holds the same body under 16 MB.
Stack allocation is a claim you can check rather than hope for. go build -gcflags=-m prints moved to heap: for anything that escapes; it prints nothing for the bookkeeping array. It does list six that escape: five are errors.As targets inside error branches, and the sixth is a bytes.Buffer used only when a type's own JSON decoding is handed query or form text. None of them is on the path a request that binds cleanly takes.
Make Errors Pay for Themselves
binder reports every failing field at once, each named by its path, such as items[2].qty or filter[team]. Those names are strings, and strings cost allocations, so they're built only when something has actually failed.
That rule is easy to break without noticing. During a later refactor we rewrote binder's JSON map loop more compactly:
// Every entry built its error name, failing or not
collectFailure(&errs, decode(dec, elem), pathPart{field: mapKeyField(k), name: mapKeyName(k)})
Go evaluates arguments before the call, so mapKeyField and mapKeyName now ran for every entry, quoting and concatenating key names nobody would read. Our six-entry map benchmark went from 35 allocations to 53. The fix restores the old shape:
if err := decode(dec, elem); err != nil {
// The entry's name is built only when it has failed.
collectFailure(&errs, err, pathPart{field: mapKeyField(k), name: mapKeyName(k)})
}
Nothing broke and every test passed. Only the benchmark caught it, which is the case for running benchmarks on every change rather than once before a release.
Don't Trust the Client With Your Buffers
Reading a body means a buffer. Sizing it from Content-Length is a single allocation of the right size, and a declared length over the limit is rejected before anything is read. But trusting Content-Length completely means a client can declare a body just under your limit, send nothing, and make you allocate the whole limit, on every connection it opens.
binder sizes the buffer from the declared length, capped at 64 KB, and grows it by doubling as data actually arrives, never past the configured limit. A body larger than 64 KB costs a few extra allocations; a lying client costs nothing. It reads at most one byte past the limit, which is enough to tell an oversized body from one exactly at the limit.
This change moved bytes far more than allocations, which is the argument for reading both columns. io.ReadAll starts at 512 bytes and doubles, so a 70-byte JSON body was costing a 512-byte buffer: B/op on the JSON case fell 79% while allocs/op fell by 2.
Two smaller wins sit next to it:
- The body is put back with one allocation instead of two. Later handlers can read it again because binder restores it as a
bytes.Readerwith aClosemethod on it, whereio.NopCloser(bytes.NewBuffer(b))would cost two. - Form bodies are parsed from the bytes already read, rather than through
Request.ParseForm— which applies its own 10 MB cap whatever your limit says, reads a body only for POST, PUT and PATCH, and reports a malformed query as if the body were at fault. A form request went from 26 allocations to 11, and from 2,616 bytes to 552.
What's Left
The query-string benchmark makes one allocation, and binding causes none of it: it's the target struct escaping to the heap when it's passed to Bind as any. Every binding library pays that one. The body benchmarks include two more for re-arming the request body between iterations, which no library can avoid, and binder pays one the others don't, to restore r.Body so later handlers can read it. Compare like with like: some of the gap between these libraries is features, not waste.
| Request | binder | Echo 4.15 | Gin 1.12 | Hand-written |
|---|---|---|---|---|
| Query, 5 fields | 527 ns, 64 B, 1 alloc | 896 ns, 544 B, 8 allocs | 1,154 ns, 608 B, 9 allocs | 360 ns, 480 B, 7 allocs |
| JSON body, 5 fields | 633 ns, 256 B, 8 allocs | 827 ns, 681 B, 8 allocs | 902 ns, 681 B, 8 allocs | 753 ns, 681 B, 8 allocs |
| Every source | 614 ns, 248 B, 6 allocs | 1,357 ns, 1,202 B, 15 allocs | 1,492 ns, 1,644 B, 21 allocs | – |
On a JSON body binder allocates the same 8 times as Echo, Gin and a hand-written encoding/json handler, but uses 256 bytes where all three use 681. The same requests through gorilla/schema, which decodes a url.Values and nothing else, cost 1,932 ns and 46 allocations on the query.
The hand-written query version is still the floor, and still faster, by about 1.5×. It returns its struct by value, so even the target doesn't allocate. That gap is roughly what reflection costs, and it's a fair price for not writing parsing code in every handler.
There is one more allocation we could take and won't. Each string field owns its value, which costs one allocation per string. Making those fields substrings of a single body string would remove it — and would mean that keeping one field keeps the whole request body alive. For a request binder, whose output often outlives the request, that trades a measurable allocation for an unbounded retention bug. It is the point where fewer allocations stops being better, and it is worth knowing where yours is.
Form bodies, at 11 allocations, and multipart, at 78, are both still above where they could be. Either could be a follow-up.
Keeping It There
Allocation work decays: someone adds a convenient fmt.Sprintf on a hot path, or a helper that takes any, and nobody notices. binder has a few guards against that:
- Benchmark tables are regenerated from clean runs, so a regression shows up as a changed number in a diff.
- Regression tests measure memory directly. Several read
runtime.MemStats.TotalAllocaround a bind and fail if it exceeds a fixed budget for that body. They exist because each one was once a real amplification: an 8 MB array body that allocated 433 MB, a deeply nestedomitemptyvalue that went quadratic in its depth at 388 MB, a million bad values that allocated 729 MB to report every one. A budget in a test is how you find out you've reintroduced it. - Differential fuzz tests keep the hand-written scanners honest against the standard library, on every CI run rather than once.
Behaviour was pinned as the work went along rather than after it: the allocation work alone added 13 test functions and the two fuzz targets above. Where it was possible, each new test was run against the old code first and had to pass there too — which is what makes it a description of binder's behaviour rather than a description of the new implementation.
Takeaways
- Measure with
-benchmem, take medians, and compare versions alternately. When something you didn't touch moves, suspect the machine. - Expect the allocation profile not to add up. The runtime packs small allocations into shared blocks, so the profiler undercounts short strings.
- An interface holding a string or a struct usually means a heap allocation. Keep hot paths typed — the costliest thing we found was a return type, not an algorithm.
- Don't parse what you won't read. Scan for the keys you need, and check the scanner against the standard library with differential fuzzing.
- Look up maps with
m[string(b)]rather than converting first. - Resolve reflection once per type and cache it.
- Keep per-request bookkeeping on the stack, check it stayed there with
-gcflags=-m, and leave failure listsniluntil something fails. - Build error strings only when there is an error, and remember that Go evaluates function arguments eagerly.
- Size buffers from what the client declares, but never trust it with more than you'd allocate anyway.
- Not every change is faster. One of ours traded 4% on a benchmark for 7 allocations, and an obvious optimisation measured worse and was reverted. Both were only knowable by measuring.
binder is open source, MIT licensed, and has no dependencies outside the standard library. The code is at github.com/uRadical/binder, the docs and full benchmark method at gobinder.dev.