Benchmarks
Festina vs. Rust, Go, and Bun
How Festina compares to Rust, Go, and Bun on a handful of
small, equivalent-logic programs, and to a browser's
<canvas> and MonoGame on 2D drawing. Not a
claim that Festina is faster than any of these languages in
general — a compiled language's real-world performance
depends heavily on what's actually being written and how
mature its optimizer/runtime is, and Festina's is young. This
exists to catch regressions and track progress over time, run
against the same few workloads on every change that plausibly
affects performance (codegen, runtime, or the standard
library), not as a marketing claim.
These same programs are also benchmarked cross-compiled
to wasm32-wasi (against C and Go, also compiled
to wasm) — see wasm.md.
Methodology #
Six programs, each implemented equivalently in Festina, Rust, Go, and Bun (source in benchmarks/), plus one comparing Festina's canvas against a browser's and MonoGame's — see Canvas at the end:
| hello | Process startup + runtime init cost — compile, print one line, exit. Directly reflects the binary-slimming work: fewer dynamically linked libraries means less for the dynamic linker to resolve before main() even runs. |
| fib | Recursive function-call overhead and raw compute throughput — naive recursive fib(32) (no memoization), ~7 million calls. Deliberately not reducible to a closed form by an optimizer (unlike a linear sum), so this actually measures generated-code quality, not the compiler's algebra. |
| loop_sum | Tight-loop / branch-free arithmetic throughput — a 100,000,000-iteration polynomial-hash accumulation (total = (total * 1000003 + i) % 1000000007), each iteration depending on the last so it can't be folded into a closed-form constant either — a plain running-sum version of this loop optimizes away entirely, running in ~2ms regardless of iteration count. |
| array_sum | Allocation-heavy throughput — 2,000,000 iterations, each building a fresh 8-element arr[int] literal (never escaping, so Festina reclaims it at that iteration's own scope-exit — see todo.md) and summing its elements into a running total. Directly exercises automatic memory management: every iteration is a genuine allocate-fill-read cycle, not just arithmetic. Each element's value depends on the previous iteration's own running total, the same closed-form-resistance trick loop_sum already uses. The hot loop lives inside a void func run(...), not bare top-level code — escape analysis only ever analyzes a function/handler's own body, never the top-level statement sequence. |
| string_concat | String-heavy throughput — 15,000 iterations of repeated concatenation (`${s}x`/s = s + "x"), s growing by one character each time. Written as the textbook O(n²) naive-concatenation pattern; Festina compiles that exact shape as an in-place append onto s's own buffer (amortized O(1) each), so its row measures that path rather than a quadratic copy. |
| char_scan | Character-by-character scan throughput — walk a ~1.7 MB buffer counting identifier runs, indexing one character at a time. The Festina version uses ascii rather than text, since a byte-indexed scan is exactly the workload ascii exists for (see api.md); Go and Rust deliberately index raw []byte/as_bytes() rather than ranging a string or using char_indices, both of which would measure UTF-8 decoding instead of scanning. |
Each language uses its own normal toolchain and optimization
settings (festina program.f -o program,
rustc -O, go build, bun
run — Bun has no separate build step, it's a JIT).
All four are checked to produce byte-identical stdout before
a run is trusted.
Every benchmark is timed with 1 untimed warmup run (page cache, dynamic linker resolution, ...) followed by 7 timed runs, keeping the minimum — the standard way to reduce OS scheduling noise without pulling in a dedicated benchmarking tool. Binary size is the compiled executable's size on disk (n/a for Bun, which ships no separate binary).
Build time gets one untimed throwaway build per toolchain before any timed one, for the same reason each program gets an untimed warmup run — without it, the first benchmark in the list would absorb the whole toolchain's own cold-start cost and report a build time several times every other program's own.
Reproduce locally:
$ python3 benchmarks/run_benchmarks.py # print results $ python3 benchmarks/run_benchmarks.py --update-doc # regenerate this file's table $ python3 benchmarks/canvas/run_canvas_benchmark.py # the canvas comparison $ python3 benchmarks/canvas/run_canvas_benchmark.py --update-doc # The MonoGame side needs a .NET SDK and, on first run, network access # to restore its NuGet package; without either it is skipped with a note # rather than failing the run. $ python3 benchmarks/http/run_http_benchmarks.py # the HTTP server comparison $ python3 benchmarks/http/run_http_benchmarks.py --update-doc # Needs `wrk` on PATH (not a project dependency -- apt/brew install wrk).
The runner skips any language toolchain not installed rather than failing — see setup.md for what each one needs.
Results #
Last run: 2026-09-10 on this machine — absolute numbers vary by hardware, relative ordering is the point.
| hello | Run time | Build time | Binary size |
|---|---|---|---|
| Festina | 1.5 ms | 93.7 ms | 1.49 MB |
| Rust | 1.7 ms | 110.2 ms | 3.77 MB |
| Go | 1.7 ms | 224.0 ms | 2.11 MB |
| Bun | 12.6 ms | n/a | n/a |
| fib(32) | Run time | Build time | Binary size |
|---|---|---|---|
| Festina | 9.1 ms | 99.7 ms | 1.49 MB |
| Rust | 9.9 ms | 118.7 ms | 3.77 MB |
| Go | 14.5 ms | 215.4 ms | 2.11 MB |
| Bun | 35.7 ms | n/a | n/a |
| loop_sum (100M) | Run time | Build time | Binary size |
|---|---|---|---|
| Festina | 526.6 ms | 105.5 ms | 1.49 MB |
| Rust | 501.3 ms | 125.5 ms | 3.77 MB |
| Go | 460.8 ms | 214.9 ms | 2.11 MB |
| Bun | 9213.8 ms | n/a | n/a |
| array_sum (2M) | Run time | Build time | Binary size |
|---|---|---|---|
| Festina | 86.4 ms | 122.4 ms | 1.49 MB |
| Rust | 86.6 ms | 159.6 ms | 3.77 MB |
| Go | 88.0 ms | 214.0 ms | 2.11 MB |
| Bun | 2674.1 ms | n/a | n/a |
| string_concat (15K) | Run time | Build time | Binary size |
|---|---|---|---|
| Festina | 1.7 ms | 109.8 ms | 1.50 MB |
| Rust | 1.7 ms | 144.0 ms | 3.77 MB |
| Go | 32.3 ms | 198.4 ms | 2.11 MB |
| Bun | 14.2 ms | n/a | n/a |
| char_scan (1.7MB) | Run time | Build time | Binary size |
|---|---|---|---|
| Festina | 13.6 ms | 166.9 ms | 1.50 MB |
| Rust | 16.3 ms | 184.8 ms | 3.77 MB |
| Go | 14.8 ms | 182.0 ms | 2.11 MB |
| Bun | 40.6 ms | n/a | n/a |
Reading these numbers #
hellois dominated by process startup, not language performance — a graphics/audio-free Festina binary (see security.md) dynamically links only libc/libm (plus libz, a transitive dependency of the statically-linked sqlite3), the same ballpark as Go's or Rust's own small dynamic dependency lists here; a graphics- or audio-using Festina program would show up slower purely from the extra shared libraries the dynamic linker has to resolve at startup (libcairo/libX11/libasoundand their own transitive dependencies).fibandloop_sumare closer to an apples-to-apples compiled-code comparison — Rust and Go both compile through mature, years-optimized backends (LLVM and Go's owngc, respectively); Festina also compiles through LLVM (see api.md) but is a much younger frontend with far less codegen-level tuning, so a gap here reflects the compiler's maturity, not a ceiling in the language design. Bun's JIT has to warm up during the run itself, which a single-shot benchmark like this doesn't isolate from the actual computation — a longer-running workload would tell a different story for Bun specifically.array_sumlands close to Rust/Go rather than behind them: the per-iterationarr[int]literal provably never escapes its own iteration, so Festina stack-allocates its header the same way a non-escaping struct local does, leaving only the growable data buffer's ownmalloc(a truly general growable buffer isn't safe to give a fixed-sizealloca). The remaining, small gap is ordinary codegen-maturity noise, not an allocation-strategy gap.string_concatis where Festina'stextownership model shows up directly. A text binding's buffer is exclusively its own, sos = `${s}x`is an assignment that is about to free the very buffer it is copying from — and the compiler treats it as what it is: an append ontos's own buffer, grown in place with a length the compiler tracks, amortized O(1) per step instead of a fresh copy of the whole string. That is the same idea Rust'sString+uses (reusing the left operand's spare capacity, likeVec), which is why the two land together. Go's+on immutable strings has no spare capacity to grow into, which is why it does the quadratic copy; Bun's V8 backend uses rope/cons-string representations internally, deferring the copy until the string is actually read, which is why it avoids the blowup despite naive-looking source. None of this is a bug in any of the four — it's exactly the kind of language/runtime difference this benchmark exists to surface.char_scanis the workloadasciiexists for: walk a ~1.7MB buffer character by character, counting identifier runs. On atextthis is quadratic — UTF-8 is variable-width, sos[i]walks from byte zero on every index — which is why the Festina version usesascii, where one byte per character puts the length in the value's own header and makes.length/s[i]/charCodeAt(i)O(1).charCodeAt(i)is emitted inline — a null check, a header load, a bounds check and a byte load, right where the expression is used — so the scan loop makes no call per character, which is what puts it level with Rust and Go rather than behind them. The Go and Rust implementations deliberately index[]byte/as_bytes()rather than ranging a string or usingchar_indices, both of which decode UTF-8 and would measure decoding instead of scanning.- The canvas comparison (below) is the one benchmark here that isn't against another language. It's against the thing a 2D game would otherwise most likely be written on: an HTML
<canvas>. Circles dominate frame cost, because Cairo tessellates every arc afresh — Festina caches one alpha mask per radius and stamps it thereafter (the same trick a glyph cache uses), which is most of why the frame stays fast. Festina also wins startup by more than an order of magnitude and wins on variance, which for a frame budget is not a footnote. - MonoGame joins the canvas comparison as a third side, and its number is the one on this page most likely to be quoted out of context. It is a GPU framework running here with no GPU, on Mesa's software rasterizer; on real hardware it would batch these 40,000 sprites into a couple of draw calls and beat everything else on this page by orders of magnitude. The row is worth having because headless rendering with no GPU is a real situation — CI, a build server, a container — and it is worth reading only with that sentence attached.
- These six are intentionally small, fast benchmarks so they can be re-run on every change worth checking, not a comprehensive suite (no concurrency, no realistic mixed workload). I/O has its own section below — see HTTP.
Canvas: Festina vs an HTML <canvas> vs MonoGame #
Last run: 2026-09-02 on this machine. Chromium 141.0.7390.37.
20,000 filled rectangles and 20,000 filled circles, fill colour changed between every shape, into an 800x600 surface. Both sides draw offscreen, both time their own draw loop with their own monotonic clock, and the browser is forced to rasterize inside the timed region. All three of those matter and all three are easy to get wrong — see run_canvas_benchmark.py, which documents what each one cost when it was measured the other way.
| Canvas | Frame (min) | Frame (median) | First frame |
|---|---|---|---|
| Festina | 7 ms | 8 ms | 21 ms |
| HTML <canvas> (Chromium/Skia) | 65 ms | 67 ms | 258 ms |
| MonoGame (software GL) | 172 ms | 188 ms | 162 ms |
llvmpipe, a software
implementation of the whole graphics pipeline. It is
therefore paying in software for vertex transform,
rasterization setup and per-pixel texture sampling that real
hardware does for free. On an actual GPU these 40,000
sprites batch into a couple of draw calls and finish in well
under a millisecond — which no CPU rasterizer on this page
can approach. What this row measures is the headless, no-GPU
case (CI, a build server, a container), and nothing else.
It is also by far the noisiest row: llvmpipe is
multithreaded and so is far more exposed to whatever else
the machine is doing than single-threaded Cairo. Consecutive
runs of the same binary measured 173, 180, 193, 285 and 498
ms. The runner launches the process five times and keeps the
best, which lands near the floor most of the time — but
treat this number as "a few hundred milliseconds," not as a
figure precise to the millisecond the way the other two rows
are.
On this workload Festina draws it 9.3x faster.
That took two changes, and finding each took measuring rather
than guessing. The first version of this benchmark had
Festina 1.4x slower, and the obvious culprit
— a fresh Cairo context per draw call — turned out to account
for 4 ms of 90. Splitting the frame by shape type found the
real one immediately: 20,000 rectangles cost 10 ms and
20,000 circles cost 76 ms, because cairo_arc +
cairo_fill tessellates the curve into Beziers
and scan-converts a general polygon every single time.
Rasterizing each radius once into an alpha mask and stamping
it thereafter — what a glyph cache does — took circles to
20 ms and the frame from 90 ms to 31 ms, leaving 11 ms of
rectangles and 20 ms of circles.
The second change noticed that neither of those needs a rasterizer at all. An opaque flat-colour rectangle at integer coordinates covers whole pixels, so its result is the colour written into each of them; an opaque circle's per-pixel coverage is the same for every circle of that radius, so Cairo rasterizes it once and the runtime blends it by hand thereafter with pixman's own 8-bit arithmetic. Every such call now writes straight into the ARGB32 pixels — no context, path, compositor dispatch or pixman call per shape — and the pixels are byte-identical to what Cairo's mask stamp produced (verified by drawing the same scene both ways, not by eye). That took 20,000 rectangles from 11 ms to 2 ms and 20,000 circles from 20 ms to 6 ms. Setting the fill colour 20,000 times is still too cheap to measure. Anything the contract does not cover — a translucent fill, a gradient, a border, a scaled or rotated canvas — still goes through Cairo exactly as before.
Two things are worth reading alongside the headline. The browser's frame time is far noisier — 65 ms at best against a 67 ms median here, and the median moves by 20+ ms between runs of this same script, while Festina's two numbers (7 and 8 ms) sit on top of each other. For a frame budget, predictability is not a footnote. And getting to the first frame differs by more than an order of magnitude in the same direction, because one side starts a process and the other starts a browser.
Both outputs were compared cell-by-cell over a 16x16 grid to confirm they drew the same scene — worst per-channel difference 0.2 out of 255. Not byte-for-byte: Cairo and Skia disagree about antialiasing on every curve, and demanding identical bytes would only prove the two rasterizers are the same program. The check has earned itself twice now: once catching a bug in this very script that left one side comparing a blank canvas, and again catching itself comparing raw RGB without accounting for alpha — Festina's own offscreen canvas starts transparent (api.md's own "a fresh or cleared canvas is transparent, not white"), so a background pixel neither side actually drew on read as black here against the browser harness's own opaque white fill, which the comparison mistook for a real rendering difference until it started compositing both sides onto the same white background first.
HTTP: Festina vs Rust vs Go vs Bun #
Four servers (source in
benchmarks/http/),
each answering the same two routes — / (a fixed
plaintext body) and /json (a small JSON body) —
load-tested with
wrk (not a project
dependency; installed separately, apt install
wrk/brew install wrk).
Equivalent logic, not equivalent idiom, the same
rule the six programs above already follow.
Festina's HTTP server (festina_runtime_http.c)
is deliberately single-threaded (one connection serviced at
a time). Rust's and Go's servers here are hand-rolled
raw-socket implementations with a single-threaded,
sequential accept loop — not hyper/net/http's
own default (multi-threaded) servers, which would be
measuring a mature framework's concurrency model against
Festina's single-threaded one rather than the same
connection-handling logic in four languages. Bun is the one
exception: it uses Bun.serve(), its own native
HTTP implementation, since there is no reason to hand-roll
sockets in a runtime that ships a fast one already (the same
"each language uses its own normal toolchain" rule the
Methodology above states).
Every response closes the connection, matched
uniformly across all four servers. Rust's and Go's
raw-socket servers close by default; Bun's server sets
Connection: close explicitly, to opt out of its
own native keep-alive (which none of the raw-socket
languages here have an equivalent of). Festina supports
HTTP/1.1 keep-alive by default (see
Keep-alive), so the
load generator itself sends the fix:
run_http_benchmarks.py's own wrk
invocation sends an explicit Connection: close
request header uniformly against all four servers —
Rust/Go/Bun ignore it (they already always close), and
Festina's own documented behavior (an explicit client
Connection: close always forces it off,
per-request) makes it close too. This keeps the comparison
to exactly connection-accept + parse + respond, for all four
languages at once, with no per-language server code needed
to special-case it.
Each wrk run: 4 threads, 50 open connections, 5
seconds, against one route at a time (a JIT-inclined runtime
like Bun gets no separate warmup here — wrk's
own 5-second window includes whatever warmup happens inside
it, the same "the timed window is the real number" approach
as the process-startup benchmarks above, since a resident
server process outlives any of them anyway).
Reproduce locally:
$ python3 benchmarks/http/run_http_benchmarks.py $ python3 benchmarks/http/run_http_benchmarks.py --update-doc $ python3 benchmarks/http/run_http_benchmarks.py --duration 10s --connections 100 --threads 8
Last run: 2026-09-01 on this machine, wrk -t4 -c50 -d5s per route.
| plaintext (/) | Requests/sec | Avg latency | Transfer/sec |
|---|---|---|---|
| Festina | 31,264 | 1.46 ms | 3.04 MB/s |
| Rust | 47,327 | 0.88 ms | 4.38 MB/s |
| Go | 26,873 | 1.68 ms | 2.49 MB/s |
| Bun | 26,618 | 1.73 ms | 3.40 MB/s |
| json (/json) | Requests/sec | Avg latency | Transfer/sec |
|---|---|---|---|
| Festina | 31,353 | 1.47 ms | 3.65 MB/s |
| Rust | 45,013 | 0.93 ms | 5.02 MB/s |
| Go | 24,138 | 1.87 ms | 2.69 MB/s |
| Bun | 26,725 | 1.75 ms | 3.92 MB/s |
Reading these numbers #
festina_http_sendcoalesces the status line and headers into a single bufferedsend()call, rather than writing each piece separately — this matters withTCP_NODELAYset (Nagle's algorithm disabled for low latency), since each separate call would otherwise become its own TCP segment. Festina clears Go on both routes and lands right around Bun's own number, with Rust's raw-socket implementation ahead of all three.- This measures connection-accept + request-parse + respond throughput under load from one client machine talking to one server process on the same machine (no network hop, no TLS) — not a claim about production capacity, the same disclaimer every other benchmark on this page already carries.
/jsonexercises more than/: Festina's route builds a struct and renders it through the same JSON-via-.toText()path every other container response already uses, not a hand-built string the way/sends one — so a gap between the two routes for Festina specifically reflects that serialization cost, not connection handling.- Rust's and Go's numbers here are not what those languages' idiomatic HTTP stacks would report — seeing "Rust is only Nx faster than Festina at HTTP" from this section should be read as "at matching, single-threaded connection handling," not as a claim about
hyper/axumornet/httpin general, which support keep-alive by default too but add a multi-threaded accept loop and years of tuning this comparison deliberately holds constant. - No WebSocket throughput benchmark exists yet —
on messagetraffic has a very different shape (persistent connections, small frequent frames) from a request/response load test, and would need its own methodology rather than reusingwrk's HTTP-request model.
HTTP: single-threaded vs. thread pool #
The section above measures Festina's single HTTP event loop
against other languages' own single-threaded raw-socket
servers — a deliberately fair, apples-to-apples comparison.
This section instead compares Festina against
itself: what a
thread pool[N] { on request(req:http) { ... } }
(a private per-thread HTTP
context) plus NAME.giveRequest(r)
actually buys a program that does real CPU-bound work per
request, the pattern
examples/threaded_http_server.f
demonstrates.
Two servers (source in
benchmarks/http_threaded/),
both answering the same two routes — / (no work,
a control) and /slow (a
closed-form-resistant polynomial-hash loop, the same
technique loop_sum.f above uses, tuned to
~2,000,000 iterations so a single request takes a few
milliseconds of real CPU time):
- single-threaded (
server_single.f) does/slow's own work directly in the one top-levelon requesthandler, on Festina's single HTTP event-loop thread — every concurrent/slowrequest queues up behind whichever one is currently computing. - thread pool[N] (
server_pool.f,N= this machine's own CPU count by default) hands every/slowrequest off to the next ofNworker threads viagiveRequest, so up toNrequests are genuinely computed in parallel, onNdifferent CPU cores, before any of them respond./is answered directly by main in both servers, unchanged — it's included to confirm the pool's own round-robin dispatch adds no meaningful overhead to a request that never needed handing off in the first place.
Each wrk run: 4 threads, 50 open connections, 5
seconds, against one route at a time — otherwise the
identical methodology the section above already uses (no
explicit Connection: close forcing here, since
both servers are the same language/runtime with the same
keep-alive behavior; there's no cross-language asymmetry to
correct for). Reproduce locally:
$ python3 benchmarks/http_threaded/run_http_threaded_benchmark.py $ python3 benchmarks/http_threaded/run_http_threaded_benchmark.py --update-doc $ python3 benchmarks/http_threaded/run_http_threaded_benchmark.py --pool-size 8 --duration 10s
Last run: 2026-09-01 on this machine (4 CPUs), wrk -t4 -c50 -d5s per route, pool size 4.
| no work, control (/) | Requests/sec | Avg latency | Transfer/sec |
|---|---|---|---|
| single-threaded | 74,673 | 0.66 ms | 8.19 MB/s |
| thread pool[4] | 49,093 | 1.03 ms | 5.38 MB/s |
| CPU-bound work (/slow) | Requests/sec | Avg latency | Transfer/sec |
|---|---|---|---|
| single-threaded | 92 | 491.28 ms | 0.01 MB/s |
| thread pool[4] | 256 | 185.33 ms | 0.03 MB/s |
/slow speedup from the pool: 2.77x (pool size 4, this machine has 4 CPUs).
Reading these numbers #
/(no work) should perform about the same on both servers — neither variant's own connection-accept/parse/respond path changed at all; only whether/slow's own CPU-bound work is serialized or parallelized did. A meaningful gap here would mean the pool's own round-robin dispatch itself is expensive, not that the pool is "working" — it shouldn't be, since/never goes throughgiveRequestin either server./slow's own speedup is bounded by real CPU core count, notN— a pool bigger than the machine's own core count just adds contention, not more genuine parallelism;--pool-sizedefaults toos.cpu_count()for exactly this reason.- Every handed-off request pays a small, real hand-off latency — a receive-only worker thread's own combined loop polls on a bounded timeout (
FESTINA_THREAD_HTTP_POLL_MS, 20ms) rather than waking instantly the way a dedicated OS thread blocked onaccept()would, so under low concurrency (one request at a time, nothing else queued) a handed-off request can be slightly slower end-to-end than the single-threaded baseline answering it directly — the pool's own advantage only shows up once there's more concurrent CPU-bound work than one thread can get through serially, which is exactly what these numbers measure. - This is a same-machine, same-process-family comparison (no network hop) measuring exactly one thing — how much a real, additional workload on the SAME hardware benefits from being spread across more than one of Festina's own execution threads — not a general claim about optimal pool sizing for a production workload, which depends heavily on how CPU-bound (vs. I/O-bound) the real work actually is.
Layered canvas: multi-threaded Festina vs a browser's Worker + OffscreenCanvas #
Last run: 2026-09-03 on this machine (4 logical CPUs). Chromium 141.0.7390.37.
Four independent layers — a sparse sky, a band of hill
texture, a band of ground texture, and a full-canvas
foreground particle scatter — 40,000 draw calls total into an
800x600 surface, the same order of magnitude as the
single-threaded canvas benchmark's own 40,000 above. Both
multi-threaded runs hand each layer to its own worker (a real
OS thread on the Festina side, a real Worker on the browser
side) and get it back with no per-pixel copy across
the boundary — an img?
(api.md)
shares its reference across postMessage instead
of cloning it, and a Worker's
transferToImageBitmap() is a genuine ownership
transfer, not a copy. Compositing the finished layers onto
one final surface IS a real per-pixel blend on both sides
(the canvas drawImage() on each, which takes an
img? layer directly on the Festina side) and
both runs time it, not just the parallel drawing — see
run_layered_canvas_benchmark.py
for the rest of what makes this comparison fair, the same
three rules the canvas benchmark's own runner already
established.
| Layered canvas | Single-threaded | 4 threads/Workers | Speedup |
|---|---|---|---|
| Festina (img?) | 11 ms (median 11 ms) | 6 ms (median 6 ms) | 1.83x |
| Browser (Skia, OffscreenCanvas) | 73 ms (median 80 ms) | 56 ms (median 57 ms) | 1.30x |
On this workload, both multi-threaded, Festina's threads draw it 9.4x faster.
Two outputs were checked, not one. Festina's multi-threaded
run was compared against its OWN single-threaded run
byte-for-byte — Cairo is deterministic, so
any difference at all would mean four threads racing to
paint four different img? buffers corrupted
something; there wasn't one. Festina's multi-threaded output
was then compared against the browser's, over the same
tolerant 16x16 grid the canvas benchmark's own runner uses
(Cairo and Skia disagree about antialiasing on every circle,
so exact bytes would only prove the two rasterizers are the
same program) — same scene both times.
Read the speedup column with the workload's own shape in
mind. The four layers are NOT equal-sized
(8,000/9,000/11,000/12,000 draws), so four threads finish in
roughly however long the heaviest layer takes, not in a
quarter of the single-threaded time — this measures what
four genuinely independent, unevenly-loaded workers buy on
real hardware, not an idealized 4x. And on the Festina side
the parallel part is now small: after the direct-fill
optimization above the 40,000 draw calls take about 4 ms on
one thread, so the serial work both runs share — clearing
the canvas and compositing four full-surface layers onto it
— is a real fraction of either number, and no amount of
threading touches it. (Before drawImage()
accepted an img? source directly, each layer
also had to be clip()-copied into a plain
img first — four 1.92 MB copies per frame, as
much time as all the drawing.)
When this benchmark was first written, Festina drew it in
84 ms single-threaded and 62 ms with four threads, and the
browser's Workers were 1.2x faster than Festina's threads.
Measuring where those 62 ms went found two things. Circles
onto an img were tessellated by Cairo on every
call, 32–40 ms per layer; they are now stamped from a cached
per-radius coverage mask, blended directly into the pixels,
and byte-identical. And the four threads were not running in
parallel at all: each painted a freshly allocated 1.92 MB
surface, and the page faults that materialize fresh memory
on first touch serialize across threads inside one process,
so four threads' worth of faults took four threads' worth of
time — every layer finished in the 11–12 ms a single thread
needed for all four. Surfaces are now faulted in when
created (one madvise on Linux), which took the
four-thread draw from 12 ms to 3.4 ms in an isolated
reproduction and is what the table above reflects.