This is the logical continuation of our benchmark methodology, this time applied to a specific metric. For details of the architecture of our Edge functions, see our pillar article on theRust architecture of Aurabase, or our comparison Wasmtime vs Wasmer for the architectural differences between the two runtimes.
The essentials
A WebAssembly cold start adds several phases (load, commit, compile or link, instantiate, first call) and two figures that do not include the same phases are not comparable, even if they display the same unit. Wasmer markets Instaboot as a direct marketing response to the subject, but its public figures are not included here without disaggregated methodology. Aurabase runs wasmtime version 43 as a production dependency for its Edge functions (verified in aura-functions/Cargo.toml), but has not published any reproducible cold start benchmarks to date. No Aurabase figures are given in this article: it is the method that is the subject.
Why the WebAssembly cold start has once again become a disputed marketing argument
The cold start has once again become an axis of commercial differentiation between WebAssembly runtimes, not just a subject of academic research. Wasmer makes this an explicit selling point with a feature called Instaboot, presented as a direct answer to the cold boot problem.
This reflex is exactly reminiscent of the dynamic already documented in our article on the backend benchmark methodology: several competing vendors display performance figures on their own product page, without always specifying the protocol that produced them. A cold start figure without a method proves nothing more than a latency figure without a method.
A cold start figure published on a marketing page, without material, without workload and without a clear definition of the starting point and the ending point of the measurement, is indistinguishable from a slogan. This applies to all runtimes cited in this article, including Aurabase on the day a figure is released.
What a cold start actually measures, and why the definition changes everything
A cold start is not a single operation: it is a sum of distinct phases, and two suppliers do not necessarily measure the same phases under the same name.
| Loading | Recovery of the .wasm module: network, disk, or already present in memory |
|---|---|
| Validation | Checking the WebAssembly bytecode structure before execution |
| Compiling or linking | JIT on the fly (Cranelift, LLVM) or linking an already precompiled artifact (AOT) |
| Instantiation | Allocation of linear memory, tables and globals, execution of a possible start function |
| First call | Processing of the request itself, sometimes included in the announced figure, sometimes excluded |
A figure that only counts the instantiation of a module already loaded and already compiled into memory will mechanically look better than a figure that includes network loading and compilation. Neither is false in itself: the problem appears when we compare them without specifying which of the two was measured.
What the academic literature shows, and why its figures do not compare with each other
Academic work published as a preprint on arXiv in recent years has measured the instantiation time of WebAssembly modules on different runtimes. The common point between these works is not a convergent figure: it is a significant difference depending on the runtime tested, the size of the module and the hardware used.
We voluntarily do not include any precise figures taken from these publications in this article. Without having thoroughly double-checked the methodology of each paper at the time of writing, republishing an isolated number would reproduce exactly the problem this article documents: a number without the context that would allow us to know what it really measures.
What this variance learns, on the other hand, is directly useful: a cold start strongly depends on the measurement context, exactly as recalled by the discipline of environment parity described in our general methodology article (identical hardware, same region, same cache state for all systems compared).
Wasmtime, Wasmer, WasmEdge: different compilation priorities
The three standalone WebAssembly runtimes most cited in this debate do not balance compilation speed and execution performance in the same way, which partly explains why their cold start figures do not compare term to term.
| Wasmtime | Cranelift compilation backend, with historically a possible precompilation route before deployment | Used in production by Aurabase for Edge functions |
|---|---|---|
| Wasmer | Has long documented several interchangeable backends, including a backend designed for compilation speed rather than runtime performance | Markets Instaboot, a reboot by snapshot of an already initialized instance |
| WasmEdge | Public positioning focused on a quick start, with its own competitive argument | Alternative runtime also active in this marketing area |
These architectures are publicly documented by the projects themselves. We have not re-verified them version by version for this article, and they do not constitute a performance ranking. They only explain why three cold start figures displayed by three different runtimes can all be accurate and yet not comparable with each other.
For the complete architectural detail between the two runtimes most often opposed in serverless discussions, see our dedicated comparison Wasmtime vs Wasmer.
What Instaboot shows, and what its product page alone doesn't prove
According to the product positioning that Wasmer communicates publicly, Instaboot restores an instance already initialized by a snapshot mechanism, rather than restarting a complete startup with each request. It is a real architectural choice that is consistent with the problem it targets.
What this article does not do, however, is repeat a performance figure displayed on the Wasmer product page. Without knowing what hardware, what workload, and what measurement protocol produced this figure, republishing it would make exactly the mistake documented above: treating a marketing number as an independent benchmark result.
A figure like “cold start less than 1 ms” has already circulated publicly, including in previous Aurabase content, without being backed by a reproducible benchmark. It is now treated internally as unsupported. The same rule applies to any figure displayed by a competing runtime, Instaboot included, as long as no disaggregated methodology accompanies it.
What Aurabase can say today about its own cold start, and what it cannot
Aurabase runs its Edge functions on Wasmtime in production, not in a pilot project. Here's exactly what the filing allows for, and where that claim ends.
The dependency is declared hard, with the features async and cranelift activated, in the production dependencies of the service, not in a dev-dependency nor a comment:
What this file does not say: no cold start figures measured according to the protocol described in our benchmark methodology exist today in the repository for this runtime. Until a dated measurement, with disaggregated percentiles, hardware and workload, has been published, no Aurabase figure should be cited as a measured characteristic of the product. For the general architecture of the platform, see our pillar article Aurabase Rust architecture. For a concrete comparison between cold start WASM and cold start container, independent of the methodological question addressed here, see our dedicated article WASM vscontainers.
How to Read a Cold Start Number Before You Believe It
Seven questions to ask any cold start figure, including ours the day we release one.
- What phases are included? Network loading, validation, compilation, instantiation, first call: a figure that only counts part of it is not comparable to a figure that counts them all.
- Was the module really “cold”? A module already in memory or disk cache does not test the same thing as a module loaded for the first time.
- JIT compilation or precompiled artifact (AOT)? The two strategies have structurally different start-up costs.
- Single digit or distribution? A best run in ten does not have the same value as a p95 in a thousand runs.
- Equipment and region specified? A figure without a hardware specification cannot be reproduced by a third party.
- Comparison at equal load and topology? Comparing a self-hosted runtime to a managed service without reporting it distorts the reading.
- Date and version of the runtime tested? An undated figure on a project that is evolving quickly means nothing after a few months.
FAQs
External sources cited: public product documentation Wasmer (Instaboot), public project documentation Wasmtime (Bytecode Alliance), consulted in preparation for this article, without independent double-checking of the performance figures they show.