This article answers a specific question — how to benchmark a backend API in a reproducible way — by documenting the protocol we will apply at Aurabase before publishing any performance figures. Not results: a method. Any measurement already published elsewhere on this site (notably on our Performancepage) which does not already rely on this protocol should be treated as unverified until further notice.
The essentials
No Aurabase performance results following a published, reproducible protocol exist to date — this article documents the methodology we will apply to produce them, not results already obtained. The repository already contains a 3-level test suite: Criterion.rs micro-benchmarks on 3 crates, k6 load tests on 8 HTTP/WebSocket scenarios, and a Python script for direct Postgres vs API comparison with percentile calculation. The complete protocol — measurement duration, percentiles rather than averages, environment isolation, version and date disclosure — is based on verified external sources: PostgreSQL, Criterion.rs, k6 (Grafana), HdrHistogram, PlanetScale and Convex. Any performance claims already published elsewhere on this site without being traced to this protocol should be considered unverified.
Why we don't publish bare figures
Convex, publisher of a competing responsive backend, has publicly distanced itself from what its technical team calls the “bar chart war” between database providers. His formula is direct: “It’s scaling theater, not scaling” — scaling theater, not real scaling (stack.convex.dev/on-competitive-benchmarks, Stack technical blog, accessed August 23, 2026).
Its central argument: a benchmark that compares two systems with different guarantees of consistency, topology or pricing model often does not test the same thing, even when it claims to do so — “the benchmark is not actually testing the same thing”. This reflex has a name in the industry: benchmarketing, publishing a figure chosen for its marketing effect rather than for its methodological rigor.
Our response is not to refuse to measure — refusing to publish a figure indefinitely would be as dishonest as publishing an unsubstantiated one. It is to first document how we would measure, with what tools and under what conditions, before claiming to have measured anything. This is also what distinguishes a useful comparison (like our Aurabase vs Appwrite comparison, which documents verifiable architectural differences) from a comparison of performance figures without a common protocol.
This is particularly important for a technical lead or a CTO who must defend a backend choice in a technical committee: a figure that cannot be traced back to a method does not survive the first, somewhat insistent question. A documented protocol is self-defending — you can show the script, the tested version, and run the test again in front of someone if necessary.
What Makes Most Backend Benchmarks Misleading
Two pitfalls come up systematically: comparing different topologies without reporting it, and measuring latency in a way that exactly hides the pauses that matter most to the user.
On the first point, PlanetScale explicitly documents its hardware parity constraint: each compared environment must run on computing resources (vCPU, RAM) equal to or greater than the reference instance, in the same cloud region (planetscale.com/benchmarks, “Telescope” methodology, accessed August 23, 2026). Without this discipline, a latency gap may simply reflect a larger machine — not a faster architecture.
The same principle applies to cache state and network topology. An instance that has just started (cold Postgres cache, empty connection pool, query plan not yet cached) responds structurally slower than an instance that has been running for an hour under stable load. A query from the same region as the database responds structurally faster than a cross-region query. Two benchmarks that specify neither are simply not comparable, even if they display identical units.
On the second point, the trap is called coordinated omission. HdrHistogram, the reference project on latency measurement created by Gil Tene, explains it this way: when a load generator waits for the response of a request before sending the next one (closed loop), a service pause automatically drops the number of requests sent during the pause — and therefore the number of high latency measurements recorded (github.com/HdrHistogram/HdrHistogram, accessed August 23, 2026). The project gives a concrete and quantified example: on a hypothetical system which samples its latency every 10 ms for 200 seconds, a single pause of 100 seconds in the middle of the test is enough to produce, without correction, a histogram where approximately 99.99% of the responses seem to fit under 1 ms - even though half of the real time has elapsed in this single pause.
A closed-loop load test that only sends a request after receiving the previous response systematically underrepresents long pauses. The p99 it displays may be better than the reality experienced by a real user — not because the system is fast, but because the measurement protocol "forgot" to send the queries while paused.
Why the average lies: p50, p95, p99
An average latency may seem excellent when one in twenty requests takes five times as long. This is precisely what the percentiles reveal and what the average structurally conceals.
Mechanically, there is nothing mysterious about a percentile: sort all measured latencies in ascending order, then take the value at the corresponding position. Out of 1000 sorted queries, p50 is the 500th value, p95 the 950th, p99 the 990th. A single abnormally slow request among 1000 is enough to make the p99 move - it is precisely its sensitivity to rare cases that makes it useful, where this same isolated request has almost no effect on the average.
Telltale sign: the text report that pgbench — the official PostgreSQL benchmark tool — displays by default gives an average and a standard deviation, not percentiles (postgresql.org/docs/current/pgbench.html, accessed August 23, 2026). Its official documentation also warns: “Never believe any test that runs for only a few seconds” – never believe a test that only runs for a few seconds, which applies as much to the duration as to the chosen metric.
k6, the load tool that we use for level 2 of our suite, solves this with thresholds expressed in percentile: the p(95)<500 syntax defines a pass/fail criterion — 95% of requests must respond within 500 ms — directly in the test configuration (grafana.com/docs/k6, consulted on August 23, 2026).
| p50 (median) | Half of queries are faster than this value | Completely hides the distribution tail |
|---|---|---|
| p95 | 1 in 20 queries is slower | Area where the first dissatisfied users appear |
| p99 | 1 in 100 queries is slower | Most sensitive to coordinated omission if the protocol is poorly designed |
The 3 benchmark levels already present in our repository
Publishing a methodology without real tools would be just another form of theater. The benchmarks/ folder in the Aurabase repository already contains a 3-level suite, inspired in its structure by the public Supabase methodology - the tools exist, the measured and dated results do not yet exist.
Level 1 — Micro-benchmarks Criterion.rs
Three crates of the Cargo workspace have dedicated CPU-bound benchmarks: aura-crypto (Argon2 hash, JWT HS256 — generation, validation and signature for PostgREST, AES-GCM encryption), aura-db-adapters (parsing filters and select in PostgREST format — eq., gte., in.(), relationship embeds), and aura-core (JSON serialization, schema_nameresolution, UUID validation).
aura-db-adapters specifically measures the cost of parsing PostgREST format queries — four cases for filters (simple_4, complex_10, or_group, in_large_50 with 50 values) and four for select (single columns, *, one relation embed, five embeds). This is the kind of cost invisible in a global load test: a regression on the analysis of a complex or.(...) filter would change almost nothing in the p95 of a little-used endpoint, but would become measurable on an endpoint with high traffic - hence the interest in isolating it in a micro-benchmark rather than relying solely on level 2.
aura-core takes a different approach: rather than measuring raw time, it measures throughput (Throughput::Bytes) on the JSON serialization and deserialization of internal NatsRequest/NatsResponse messages exchanged between the gateway and the services — with three realistic payload sizes (a minimal request, a request with a JSON body nested, a 50-line list response).
Criterion.rs doesn't just time a loop. It first runs a warm-up phase to fill the CPU/OS caches, detects outliers with a modified version of Tukey's method (without excluding them from the dataset), calculates confidence intervals by bootstrapping on a large number of resampled samples, and detects performance regressions between two runs by Student's statistical test, with a configurable noise threshold — typically ±1% — to ignore variations that are not statistically significant (bheisler.github.io/criterion.rs/book/analysis.html, accessed August 23, 2026).
Each Criterion run generates a detailed HTML report in target/criterion/ — distributions, regression graphs, comparison to previous run. It is this relationship, not just a terminal line, that a serious methodology must make it possible to regenerate.
Level 2 — k6 load testing
Eight k6 scripts cover the gateway on the data plane side: health (latency baseline), auth-flow (register → login → refresh → logout), crud-read and crud-write, storage (upload/download), realtime-ws, breakpoint (load increase until failure) and supabase-compare. Seven are wired to a dedicated Makefile target — supabase-compare.js exists in the repository but does not yet have a target, a state of affairs that this article documents as is rather than disguising it.
Shared configuration defines thresholds per operation type. These are pass/fail criteria that the test checks each time it runs — not results already measured:
| Reading (GET) | p95 < 500 ms · p99 < 1000 ms | Config k6 (benchmarks/k6/lib/config.js) |
|---|---|---|
| Writing (POST/PATCH) | p95 < 300 ms · p99 < 1000 ms | Config k6 |
| Auth (login/refresh) | p95 < 300 ms · p99 < 1000 ms | Config k6 |
| Storage (upload/download) | p95 < 500 ms · p99 < 2000 ms | Config k6 |
| Error rate, all scenarios | < 1 % | Config k6 |
The README.md in the benchmarks/ folder documents a reading threshold of p95 at < 200ms ("Supabase SLO"), while the threshold actually applied in benchmarks/k6/lib/config.js — the one the test runs — is p(95)<500. The two files are derived from each other. This is a concrete example, found while reading the source code for this article, of why a protocol should have a single versioned source of truth rather than being documented in two places: without it, even a team that tries to be rigorous ends up publishing conflicting thresholds.
Level 3 — Direct PostgreSQL vs API comparison
A Python script (direct_vs_api.py) measures the actual overhead of the gateway + service layer by comparing direct psycopg2 requests to HTTP calls on the same operation — list, one-time read by id, filtered and sorted read. Each measurement follows a warm-up of 10 iterations before the timed loop, then calculates average, p50, p95, p99 and a throughput in operations per second.
A second script (aurabase_vs_supabase.py) applies the same warmup and percentile calculation logic to a head-to-head comparison with a local Supabase instance (Supabase CLI, localhost:54321 by default) — same machine, same local network for both, exactly the environment parity discipline that PlanetScale documents for its own comparisons.
An orchestration script (collect_baseline.sh, target bench-baseline of Makefile) connects the three levels — Criterion on the 3 crates, a subset of the k6 scenarios (health and crud-read today, not all 8 yet), then the Python comparison — and writes logs, JSON and Criterion HTML reports to a folder unique timestamped: benchmarks/results/AAAAMMJJ_HHMMSS/. This is exactly the reflex of dated disclosure, in a single reproducible run, that the following section formalizes in a complete protocol.
The protocol we will apply before publishing a figure
Eight commitments, each anchored in a practice already documented by a recognized third-party tool or project — not invented for the occasion.
- Preheating separate from measurement. Criterion.rs fills CPU/OS caches before timing;
pgbenchexplicitly recommends never believing a run lasting just a few seconds. - Fixed duration, not a fixed number of iterations. A load needs time to converge — this is the role of
stagesk6 and the-Tflag ofpgbench. - Percentiles, never just the average — and active vigilance on coordinated omission if the load generator operates in a closed loop.
- Environment documented in detail: git commit of the service tested, version of PostgreSQL, hardware specification, version of the load tool. PlanetScale documents its exact TPCC parameters (
TABLES=20,SCALE=250, ~500 GB) for precisely this reason — without these details, no one can reproduce a run. - Timestamped and versioned results, never a single number engraved on a marketing page without a date. The current tooling is already written in a dated file; it will be necessary to extend this reflex to any publicly published measurement, with the hosting region documented like any other environmental variable (see our guide on EU hosting sovereignty, relevant as soon as a figure depends on a given region).
- Scripts and raw data published alongside the aggregate result, not just a final average. PlanetScale even invites readers to report a methodological error on a dedicated address – a posture that we find healthy and that we want to resume.
- Advertised throughput alongside latency, not just one or the other. A system can have excellent latency at low load and collapse in throughput as concurrency increases — that's exactly what our k6 suite's
breakpointscenario (scale-to-crash) is designed to reveal, and what Criterion'sThroughput::Bytesmicrobenchmark measurement captures at the function level. - Significant gap before announcing an improvement. A variation of a few percent between two runs may be measurement noise rather than a real gain — Criterion.rs calculates a probability that the observed difference is due to chance before qualifying it as regression or improvement. An isolated figure, without this verification, is just a statistical anecdote.
What we won't do
This list counts as much as the positive protocol above.
- Comparing different topologies (self-hosted vs managed, cold vs pre-heated instance) without explicitly reporting it.
- Retain the best run out of ten without mentioning the other nine.
- Publish a figure without a date, without a service version, without a reproduction script.
- Republish an existing marketing figure as long as it is not traced back to this protocol.
- Comparing ourselves to a competitor on a raw performance figure if that competitor does not publish its own methodology in an equivalent manner — a figure versus silence is not a comparison, it's a slogan.
A figure like “cold start less than 1 ms” was circulated without being backed by a reproducible benchmark. It is now treated internally as unsupported and should not be read as a measured characteristic of the product until any dated measurement, with published methodology, confirms it. This is precisely the kind of claim that this protocol exists to prevent from repeating.
The minimal protocol for benchmarking any backend
This protocol does not depend on any specific Aurabase tool — you can apply it to your own API today.
- Set the load before the tool: read-only, write, realistic mix for your application — not a generic ratio copied from another project.
- Explicitly separate the preheating phase from the measuring phase.
- Run the test long enough — minutes, not seconds.
- Measure in percentiles (p50/p95/p99), never on average alone.
- Verify that your load generator is not in closed loop, or correct the coordination omission in the analysis.
- Isolate the environment under test — no noisy neighbors, no competing background tasks.
- Publish the tested version, date, hardware spec, and script — not just the final result.
On a bare Postgres base, this protocol takes one command pgbench — 20 concurrent clients distributed over 4 threads, for 5 minutes, with a progress report every 10 seconds:
Reference tools, by level
Five tools, each suited to a different level of the stack—none replaces the others.
| Microphone (function) | Criterion.rs | Pure CPU, bootstrap statistics |
|---|---|---|
| SQL Query | pgbench | TPC-B-like transaction, tps and latency |
| HTTP/WS load | k6 (Grafana) | Percentiles, pass/fail thresholds |
| OLTP at scale | sysbench + TPCC (Telescope methodology) | QPS, cost per performance |
| Measurement correction | HdrHistogram | Compensates for coordinated omission |