f Festina / Docs

Benchmarks

Festina vs. Rust, Go, and Bun

How Festina compares to Rust, Go, and Bun on a handful of small, equivalent-logic programs, and to a browser's <canvas> and MonoGame on 2D drawing. Not a claim that Festina is faster than any of these languages in general — a compiled language's real-world performance depends heavily on what's actually being written and how mature its optimizer/runtime is, and Festina's is young. This exists to catch regressions and track progress over time, run against the same few workloads on every change that plausibly affects performance (codegen, runtime, or the standard library), not as a marketing claim.

These same programs are also benchmarked cross-compiled to wasm32-wasi (against C and Go, also compiled to wasm) — see wasm.md.

Methodology #

Six programs, each implemented equivalently in Festina, Rust, Go, and Bun (source in benchmarks/), plus one comparing Festina's canvas against a browser's and MonoGame's — see Canvas at the end:

helloProcess startup + runtime init cost — compile, print one line, exit. Directly reflects the binary-slimming work: fewer dynamically linked libraries means less for the dynamic linker to resolve before main() even runs.
fibRecursive function-call overhead and raw compute throughput — naive recursive fib(32) (no memoization), ~7 million calls. Deliberately not reducible to a closed form by an optimizer (unlike a linear sum), so this actually measures generated-code quality, not the compiler's algebra.
loop_sumTight-loop / branch-free arithmetic throughput — a 100,000,000-iteration polynomial-hash accumulation (total = (total * 1000003 + i) % 1000000007), each iteration depending on the last so it can't be folded into a closed-form constant either — a plain running-sum version of this loop optimizes away entirely, running in ~2ms regardless of iteration count.
array_sumAllocation-heavy throughput — 2,000,000 iterations, each building a fresh 8-element arr[int] literal (never escaping, so Festina reclaims it at that iteration's own scope-exit — see todo.md) and summing its elements into a running total. Directly exercises automatic memory management: every iteration is a genuine allocate-fill-read cycle, not just arithmetic. Each element's value depends on the previous iteration's own running total, the same closed-form-resistance trick loop_sum already uses. The hot loop lives inside a void func run(...), not bare top-level code — escape analysis only ever analyzes a function/handler's own body, never the top-level statement sequence.
string_concatString-heavy throughput — 15,000 iterations of repeated concatenation (`${s}x`/s = s + "x"), s growing by one character each time. Written as the textbook O(n²) naive-concatenation pattern; Festina compiles that exact shape as an in-place append onto s's own buffer (amortized O(1) each), so its row measures that path rather than a quadratic copy.
char_scanCharacter-by-character scan throughput — walk a ~1.7 MB buffer counting identifier runs, indexing one character at a time. The Festina version uses ascii rather than text, since a byte-indexed scan is exactly the workload ascii exists for (see api.md); Go and Rust deliberately index raw []byte/as_bytes() rather than ranging a string or using char_indices, both of which would measure UTF-8 decoding instead of scanning.

Each language uses its own normal toolchain and optimization settings (festina program.f -o program, rustc -O, go build, bun run — Bun has no separate build step, it's a JIT). All four are checked to produce byte-identical stdout before a run is trusted.

Every benchmark is timed with 1 untimed warmup run (page cache, dynamic linker resolution, ...) followed by 7 timed runs, keeping the minimum — the standard way to reduce OS scheduling noise without pulling in a dedicated benchmarking tool. Binary size is the compiled executable's size on disk (n/a for Bun, which ships no separate binary).

Build time gets one untimed throwaway build per toolchain before any timed one, for the same reason each program gets an untimed warmup run — without it, the first benchmark in the list would absorb the whole toolchain's own cold-start cost and report a build time several times every other program's own.

Reproduce locally:

$ python3 benchmarks/run_benchmarks.py               # print results
$ python3 benchmarks/run_benchmarks.py --update-doc   # regenerate this file's table

$ python3 benchmarks/canvas/run_canvas_benchmark.py             # the canvas comparison
$ python3 benchmarks/canvas/run_canvas_benchmark.py --update-doc

# The MonoGame side needs a .NET SDK and, on first run, network access
# to restore its NuGet package; without either it is skipped with a note
# rather than failing the run.

$ python3 benchmarks/http/run_http_benchmarks.py                # the HTTP server comparison
$ python3 benchmarks/http/run_http_benchmarks.py --update-doc

# Needs `wrk` on PATH (not a project dependency -- apt/brew install wrk).

The runner skips any language toolchain not installed rather than failing — see setup.md for what each one needs.

Results #

Last run: 2026-09-10 on this machine — absolute numbers vary by hardware, relative ordering is the point.

helloRun timeBuild timeBinary size
Festina1.5 ms93.7 ms1.49 MB
Rust1.7 ms110.2 ms3.77 MB
Go1.7 ms224.0 ms2.11 MB
Bun12.6 msn/an/a
fib(32)Run timeBuild timeBinary size
Festina9.1 ms99.7 ms1.49 MB
Rust9.9 ms118.7 ms3.77 MB
Go14.5 ms215.4 ms2.11 MB
Bun35.7 msn/an/a
loop_sum (100M)Run timeBuild timeBinary size
Festina526.6 ms105.5 ms1.49 MB
Rust501.3 ms125.5 ms3.77 MB
Go460.8 ms214.9 ms2.11 MB
Bun9213.8 msn/an/a
array_sum (2M)Run timeBuild timeBinary size
Festina86.4 ms122.4 ms1.49 MB
Rust86.6 ms159.6 ms3.77 MB
Go88.0 ms214.0 ms2.11 MB
Bun2674.1 msn/an/a
string_concat (15K)Run timeBuild timeBinary size
Festina1.7 ms109.8 ms1.50 MB
Rust1.7 ms144.0 ms3.77 MB
Go32.3 ms198.4 ms2.11 MB
Bun14.2 msn/an/a
char_scan (1.7MB)Run timeBuild timeBinary size
Festina13.6 ms166.9 ms1.50 MB
Rust16.3 ms184.8 ms3.77 MB
Go14.8 ms182.0 ms2.11 MB
Bun40.6 msn/an/a

Reading these numbers #

  • hello is dominated by process startup, not language performance — a graphics/audio-free Festina binary (see security.md) dynamically links only libc/libm (plus libz, a transitive dependency of the statically-linked sqlite3), the same ballpark as Go's or Rust's own small dynamic dependency lists here; a graphics- or audio-using Festina program would show up slower purely from the extra shared libraries the dynamic linker has to resolve at startup (libcairo/libX11/libasound and their own transitive dependencies).
  • fib and loop_sum are closer to an apples-to-apples compiled-code comparison — Rust and Go both compile through mature, years-optimized backends (LLVM and Go's own gc, respectively); Festina also compiles through LLVM (see api.md) but is a much younger frontend with far less codegen-level tuning, so a gap here reflects the compiler's maturity, not a ceiling in the language design. Bun's JIT has to warm up during the run itself, which a single-shot benchmark like this doesn't isolate from the actual computation — a longer-running workload would tell a different story for Bun specifically.
  • array_sum lands close to Rust/Go rather than behind them: the per-iteration arr[int] literal provably never escapes its own iteration, so Festina stack-allocates its header the same way a non-escaping struct local does, leaving only the growable data buffer's own malloc (a truly general growable buffer isn't safe to give a fixed-size alloca). The remaining, small gap is ordinary codegen-maturity noise, not an allocation-strategy gap.
  • string_concat is where Festina's text ownership model shows up directly. A text binding's buffer is exclusively its own, so s = `${s}x` is an assignment that is about to free the very buffer it is copying from — and the compiler treats it as what it is: an append onto s's own buffer, grown in place with a length the compiler tracks, amortized O(1) per step instead of a fresh copy of the whole string. That is the same idea Rust's String + uses (reusing the left operand's spare capacity, like Vec), which is why the two land together. Go's + on immutable strings has no spare capacity to grow into, which is why it does the quadratic copy; Bun's V8 backend uses rope/cons-string representations internally, deferring the copy until the string is actually read, which is why it avoids the blowup despite naive-looking source. None of this is a bug in any of the four — it's exactly the kind of language/runtime difference this benchmark exists to surface.
  • char_scan is the workload ascii exists for: walk a ~1.7MB buffer character by character, counting identifier runs. On a text this is quadratic — UTF-8 is variable-width, so s[i] walks from byte zero on every index — which is why the Festina version uses ascii, where one byte per character puts the length in the value's own header and makes .length/s[i]/charCodeAt(i) O(1). charCodeAt(i) is emitted inline — a null check, a header load, a bounds check and a byte load, right where the expression is used — so the scan loop makes no call per character, which is what puts it level with Rust and Go rather than behind them. The Go and Rust implementations deliberately index []byte/as_bytes() rather than ranging a string or using char_indices, both of which decode UTF-8 and would measure decoding instead of scanning.
  • The canvas comparison (below) is the one benchmark here that isn't against another language. It's against the thing a 2D game would otherwise most likely be written on: an HTML <canvas>. Circles dominate frame cost, because Cairo tessellates every arc afresh — Festina caches one alpha mask per radius and stamps it thereafter (the same trick a glyph cache uses), which is most of why the frame stays fast. Festina also wins startup by more than an order of magnitude and wins on variance, which for a frame budget is not a footnote.
  • MonoGame joins the canvas comparison as a third side, and its number is the one on this page most likely to be quoted out of context. It is a GPU framework running here with no GPU, on Mesa's software rasterizer; on real hardware it would batch these 40,000 sprites into a couple of draw calls and beat everything else on this page by orders of magnitude. The row is worth having because headless rendering with no GPU is a real situation — CI, a build server, a container — and it is worth reading only with that sentence attached.
  • These six are intentionally small, fast benchmarks so they can be re-run on every change worth checking, not a comprehensive suite (no concurrency, no realistic mixed workload). I/O has its own section below — see HTTP.

Canvas: Festina vs an HTML <canvas> vs MonoGame #

Last run: 2026-09-02 on this machine. Chromium 141.0.7390.37.

20,000 filled rectangles and 20,000 filled circles, fill colour changed between every shape, into an 800x600 surface. Both sides draw offscreen, both time their own draw loop with their own monotonic clock, and the browser is forced to rasterize inside the timed region. All three of those matter and all three are easy to get wrong — see run_canvas_benchmark.py, which documents what each one cost when it was measured the other way.

CanvasFrame (min)Frame (median)First frame
Festina7 ms8 ms21 ms
HTML <canvas> (Chromium/Skia)65 ms67 ms258 ms
MonoGame (software GL)172 ms188 ms162 ms
Read the MonoGame row with its caveat. MonoGame is a GPU framework, and this machine has no GPU — its GL context is Mesa's llvmpipe, a software implementation of the whole graphics pipeline. It is therefore paying in software for vertex transform, rasterization setup and per-pixel texture sampling that real hardware does for free. On an actual GPU these 40,000 sprites batch into a couple of draw calls and finish in well under a millisecond — which no CPU rasterizer on this page can approach. What this row measures is the headless, no-GPU case (CI, a build server, a container), and nothing else. It is also by far the noisiest row: llvmpipe is multithreaded and so is far more exposed to whatever else the machine is doing than single-threaded Cairo. Consecutive runs of the same binary measured 173, 180, 193, 285 and 498 ms. The runner launches the process five times and keeps the best, which lands near the floor most of the time — but treat this number as "a few hundred milliseconds," not as a figure precise to the millisecond the way the other two rows are.

On this workload Festina draws it 9.3x faster.

That took two changes, and finding each took measuring rather than guessing. The first version of this benchmark had Festina 1.4x slower, and the obvious culprit — a fresh Cairo context per draw call — turned out to account for 4 ms of 90. Splitting the frame by shape type found the real one immediately: 20,000 rectangles cost 10 ms and 20,000 circles cost 76 ms, because cairo_arc + cairo_fill tessellates the curve into Beziers and scan-converts a general polygon every single time. Rasterizing each radius once into an alpha mask and stamping it thereafter — what a glyph cache does — took circles to 20 ms and the frame from 90 ms to 31 ms, leaving 11 ms of rectangles and 20 ms of circles.

The second change noticed that neither of those needs a rasterizer at all. An opaque flat-colour rectangle at integer coordinates covers whole pixels, so its result is the colour written into each of them; an opaque circle's per-pixel coverage is the same for every circle of that radius, so Cairo rasterizes it once and the runtime blends it by hand thereafter with pixman's own 8-bit arithmetic. Every such call now writes straight into the ARGB32 pixels — no context, path, compositor dispatch or pixman call per shape — and the pixels are byte-identical to what Cairo's mask stamp produced (verified by drawing the same scene both ways, not by eye). That took 20,000 rectangles from 11 ms to 2 ms and 20,000 circles from 20 ms to 6 ms. Setting the fill colour 20,000 times is still too cheap to measure. Anything the contract does not cover — a translucent fill, a gradient, a border, a scaled or rotated canvas — still goes through Cairo exactly as before.

Two things are worth reading alongside the headline. The browser's frame time is far noisier — 65 ms at best against a 67 ms median here, and the median moves by 20+ ms between runs of this same script, while Festina's two numbers (7 and 8 ms) sit on top of each other. For a frame budget, predictability is not a footnote. And getting to the first frame differs by more than an order of magnitude in the same direction, because one side starts a process and the other starts a browser.

Both outputs were compared cell-by-cell over a 16x16 grid to confirm they drew the same scene — worst per-channel difference 0.2 out of 255. Not byte-for-byte: Cairo and Skia disagree about antialiasing on every curve, and demanding identical bytes would only prove the two rasterizers are the same program. The check has earned itself twice now: once catching a bug in this very script that left one side comparing a blank canvas, and again catching itself comparing raw RGB without accounting for alpha — Festina's own offscreen canvas starts transparent (api.md's own "a fresh or cleared canvas is transparent, not white"), so a background pixel neither side actually drew on read as black here against the browser harness's own opaque white fill, which the comparison mistook for a real rendering difference until it started compositing both sides onto the same white background first.

HTTP: Festina vs Rust vs Go vs Bun #

Four servers (source in benchmarks/http/), each answering the same two routes — / (a fixed plaintext body) and /json (a small JSON body) — load-tested with wrk (not a project dependency; installed separately, apt install wrk/brew install wrk).

Equivalent logic, not equivalent idiom, the same rule the six programs above already follow. Festina's HTTP server (festina_runtime_http.c) is deliberately single-threaded (one connection serviced at a time). Rust's and Go's servers here are hand-rolled raw-socket implementations with a single-threaded, sequential accept loop — not hyper/net/http's own default (multi-threaded) servers, which would be measuring a mature framework's concurrency model against Festina's single-threaded one rather than the same connection-handling logic in four languages. Bun is the one exception: it uses Bun.serve(), its own native HTTP implementation, since there is no reason to hand-roll sockets in a runtime that ships a fast one already (the same "each language uses its own normal toolchain" rule the Methodology above states).

Every response closes the connection, matched uniformly across all four servers. Rust's and Go's raw-socket servers close by default; Bun's server sets Connection: close explicitly, to opt out of its own native keep-alive (which none of the raw-socket languages here have an equivalent of). Festina supports HTTP/1.1 keep-alive by default (see Keep-alive), so the load generator itself sends the fix: run_http_benchmarks.py's own wrk invocation sends an explicit Connection: close request header uniformly against all four servers — Rust/Go/Bun ignore it (they already always close), and Festina's own documented behavior (an explicit client Connection: close always forces it off, per-request) makes it close too. This keeps the comparison to exactly connection-accept + parse + respond, for all four languages at once, with no per-language server code needed to special-case it.

Each wrk run: 4 threads, 50 open connections, 5 seconds, against one route at a time (a JIT-inclined runtime like Bun gets no separate warmup here — wrk's own 5-second window includes whatever warmup happens inside it, the same "the timed window is the real number" approach as the process-startup benchmarks above, since a resident server process outlives any of them anyway).

Reproduce locally:

$ python3 benchmarks/http/run_http_benchmarks.py
$ python3 benchmarks/http/run_http_benchmarks.py --update-doc
$ python3 benchmarks/http/run_http_benchmarks.py --duration 10s --connections 100 --threads 8

Last run: 2026-09-01 on this machine, wrk -t4 -c50 -d5s per route.

plaintext (/)Requests/secAvg latencyTransfer/sec
Festina31,2641.46 ms3.04 MB/s
Rust47,3270.88 ms4.38 MB/s
Go26,8731.68 ms2.49 MB/s
Bun26,6181.73 ms3.40 MB/s
json (/json)Requests/secAvg latencyTransfer/sec
Festina31,3531.47 ms3.65 MB/s
Rust45,0130.93 ms5.02 MB/s
Go24,1381.87 ms2.69 MB/s
Bun26,7251.75 ms3.92 MB/s

Reading these numbers #

  • festina_http_send coalesces the status line and headers into a single buffered send() call, rather than writing each piece separately — this matters with TCP_NODELAY set (Nagle's algorithm disabled for low latency), since each separate call would otherwise become its own TCP segment. Festina clears Go on both routes and lands right around Bun's own number, with Rust's raw-socket implementation ahead of all three.
  • This measures connection-accept + request-parse + respond throughput under load from one client machine talking to one server process on the same machine (no network hop, no TLS) — not a claim about production capacity, the same disclaimer every other benchmark on this page already carries.
  • /json exercises more than /: Festina's route builds a struct and renders it through the same JSON-via-.toText() path every other container response already uses, not a hand-built string the way / sends one — so a gap between the two routes for Festina specifically reflects that serialization cost, not connection handling.
  • Rust's and Go's numbers here are not what those languages' idiomatic HTTP stacks would report — seeing "Rust is only Nx faster than Festina at HTTP" from this section should be read as "at matching, single-threaded connection handling," not as a claim about hyper/axum or net/http in general, which support keep-alive by default too but add a multi-threaded accept loop and years of tuning this comparison deliberately holds constant.
  • No WebSocket throughput benchmark exists yet — on message traffic has a very different shape (persistent connections, small frequent frames) from a request/response load test, and would need its own methodology rather than reusing wrk's HTTP-request model.

HTTP: single-threaded vs. thread pool #

The section above measures Festina's single HTTP event loop against other languages' own single-threaded raw-socket servers — a deliberately fair, apples-to-apples comparison. This section instead compares Festina against itself: what a thread pool[N] { on request(req:http) { ... } } (a private per-thread HTTP context) plus NAME.giveRequest(r) actually buys a program that does real CPU-bound work per request, the pattern examples/threaded_http_server.f demonstrates.

Two servers (source in benchmarks/http_threaded/), both answering the same two routes — / (no work, a control) and /slow (a closed-form-resistant polynomial-hash loop, the same technique loop_sum.f above uses, tuned to ~2,000,000 iterations so a single request takes a few milliseconds of real CPU time):

  • single-threaded (server_single.f) does /slow's own work directly in the one top-level on request handler, on Festina's single HTTP event-loop thread — every concurrent /slow request queues up behind whichever one is currently computing.
  • thread pool[N] (server_pool.f, N = this machine's own CPU count by default) hands every /slow request off to the next of N worker threads via giveRequest, so up to N requests are genuinely computed in parallel, on N different CPU cores, before any of them respond. / is answered directly by main in both servers, unchanged — it's included to confirm the pool's own round-robin dispatch adds no meaningful overhead to a request that never needed handing off in the first place.

Each wrk run: 4 threads, 50 open connections, 5 seconds, against one route at a time — otherwise the identical methodology the section above already uses (no explicit Connection: close forcing here, since both servers are the same language/runtime with the same keep-alive behavior; there's no cross-language asymmetry to correct for). Reproduce locally:

$ python3 benchmarks/http_threaded/run_http_threaded_benchmark.py
$ python3 benchmarks/http_threaded/run_http_threaded_benchmark.py --update-doc
$ python3 benchmarks/http_threaded/run_http_threaded_benchmark.py --pool-size 8 --duration 10s

Last run: 2026-09-01 on this machine (4 CPUs), wrk -t4 -c50 -d5s per route, pool size 4.

no work, control (/)Requests/secAvg latencyTransfer/sec
single-threaded74,6730.66 ms8.19 MB/s
thread pool[4]49,0931.03 ms5.38 MB/s
CPU-bound work (/slow)Requests/secAvg latencyTransfer/sec
single-threaded92491.28 ms0.01 MB/s
thread pool[4]256185.33 ms0.03 MB/s

/slow speedup from the pool: 2.77x (pool size 4, this machine has 4 CPUs).

Reading these numbers #

  • / (no work) should perform about the same on both servers — neither variant's own connection-accept/parse/respond path changed at all; only whether /slow's own CPU-bound work is serialized or parallelized did. A meaningful gap here would mean the pool's own round-robin dispatch itself is expensive, not that the pool is "working" — it shouldn't be, since / never goes through giveRequest in either server.
  • /slow's own speedup is bounded by real CPU core count, not N — a pool bigger than the machine's own core count just adds contention, not more genuine parallelism; --pool-size defaults to os.cpu_count() for exactly this reason.
  • Every handed-off request pays a small, real hand-off latency — a receive-only worker thread's own combined loop polls on a bounded timeout (FESTINA_THREAD_HTTP_POLL_MS, 20ms) rather than waking instantly the way a dedicated OS thread blocked on accept() would, so under low concurrency (one request at a time, nothing else queued) a handed-off request can be slightly slower end-to-end than the single-threaded baseline answering it directly — the pool's own advantage only shows up once there's more concurrent CPU-bound work than one thread can get through serially, which is exactly what these numbers measure.
  • This is a same-machine, same-process-family comparison (no network hop) measuring exactly one thing — how much a real, additional workload on the SAME hardware benefits from being spread across more than one of Festina's own execution threads — not a general claim about optimal pool sizing for a production workload, which depends heavily on how CPU-bound (vs. I/O-bound) the real work actually is.

Layered canvas: multi-threaded Festina vs a browser's Worker + OffscreenCanvas #

Last run: 2026-09-03 on this machine (4 logical CPUs). Chromium 141.0.7390.37.

Four independent layers — a sparse sky, a band of hill texture, a band of ground texture, and a full-canvas foreground particle scatter — 40,000 draw calls total into an 800x600 surface, the same order of magnitude as the single-threaded canvas benchmark's own 40,000 above. Both multi-threaded runs hand each layer to its own worker (a real OS thread on the Festina side, a real Worker on the browser side) and get it back with no per-pixel copy across the boundary — an img? (api.md) shares its reference across postMessage instead of cloning it, and a Worker's transferToImageBitmap() is a genuine ownership transfer, not a copy. Compositing the finished layers onto one final surface IS a real per-pixel blend on both sides (the canvas drawImage() on each, which takes an img? layer directly on the Festina side) and both runs time it, not just the parallel drawing — see run_layered_canvas_benchmark.py for the rest of what makes this comparison fair, the same three rules the canvas benchmark's own runner already established.

Layered canvasSingle-threaded4 threads/WorkersSpeedup
Festina (img?)11 ms (median 11 ms)6 ms (median 6 ms)1.83x
Browser (Skia, OffscreenCanvas)73 ms (median 80 ms)56 ms (median 57 ms)1.30x

On this workload, both multi-threaded, Festina's threads draw it 9.4x faster.

Two outputs were checked, not one. Festina's multi-threaded run was compared against its OWN single-threaded run byte-for-byte — Cairo is deterministic, so any difference at all would mean four threads racing to paint four different img? buffers corrupted something; there wasn't one. Festina's multi-threaded output was then compared against the browser's, over the same tolerant 16x16 grid the canvas benchmark's own runner uses (Cairo and Skia disagree about antialiasing on every circle, so exact bytes would only prove the two rasterizers are the same program) — same scene both times.

Read the speedup column with the workload's own shape in mind. The four layers are NOT equal-sized (8,000/9,000/11,000/12,000 draws), so four threads finish in roughly however long the heaviest layer takes, not in a quarter of the single-threaded time — this measures what four genuinely independent, unevenly-loaded workers buy on real hardware, not an idealized 4x. And on the Festina side the parallel part is now small: after the direct-fill optimization above the 40,000 draw calls take about 4 ms on one thread, so the serial work both runs share — clearing the canvas and compositing four full-surface layers onto it — is a real fraction of either number, and no amount of threading touches it. (Before drawImage() accepted an img? source directly, each layer also had to be clip()-copied into a plain img first — four 1.92 MB copies per frame, as much time as all the drawing.)

When this benchmark was first written, Festina drew it in 84 ms single-threaded and 62 ms with four threads, and the browser's Workers were 1.2x faster than Festina's threads. Measuring where those 62 ms went found two things. Circles onto an img were tessellated by Cairo on every call, 32–40 ms per layer; they are now stamped from a cached per-radius coverage mask, blended directly into the pixels, and byte-identical. And the four threads were not running in parallel at all: each painted a freshly allocated 1.92 MB surface, and the page faults that materialize fresh memory on first touch serialize across threads inside one process, so four threads' worth of faults took four threads' worth of time — every layer finished in the 11–12 ms a single thread needed for all four. Surfaces are now faulted in when created (one madvise on Linux), which took the four-thread draw from 12 ms to 3.4 ms in an isolated reproduction and is what the table above reflects.