PRODSovereign European BaaS platformOpen Dashboard →

Performance · 15 min read

How we benchmark a backend: a repeatable methodology

Affane Daylami · Fondateur · July 8, 2026

Back to blog

A benchmark figure without a method proves nothing. “p95 under X ms”, “cold start less than Y ms” — anyone can write that on a marketing page. What proves something is the method: the equipment used, the duration of the test, the measurement protocol, and the possibility for a third party to reproduce it identically.

This English text was generated automatically from the French original and has not been reviewed yet.

This article answers a specific question — how to benchmark a backend API in a reproducible way — by documenting the protocol we will apply at Aurabase before publishing any performance figures. Not results: a method. Any measurement already published elsewhere on this site (notably on our Performancepage) which does not already rely on this protocol should be treated as unverified until further notice.

The essentials

No Aurabase performance results following a published, reproducible protocol exist to date — this article documents the methodology we will apply to produce them, not results already obtained. The repository already contains a 3-level test suite: Criterion.rs micro-benchmarks on 3 crates, k6 load tests on 8 HTTP/WebSocket scenarios, and a Python script for direct Postgres vs API comparison with percentile calculation. The complete protocol — measurement duration, percentiles rather than averages, environment isolation, version and date disclosure — is based on verified external sources: PostgreSQL, Criterion.rs, k6 (Grafana), HdrHistogram, PlanetScale and Convex. Any performance claims already published elsewhere on this site without being traced to this protocol should be considered unverified.

#
Editorial stance

Why we don't publish bare figures

Convex, publisher of a competing responsive backend, has publicly distanced itself from what its technical team calls the “bar chart war” between database providers. His formula is direct: “It’s scaling theater, not scaling” — scaling theater, not real scaling (stack.convex.dev/on-competitive-benchmarks, Stack technical blog, accessed August 23, 2026).

Its central argument: a benchmark that compares two systems with different guarantees of consistency, topology or pricing model often does not test the same thing, even when it claims to do so — “the benchmark is not actually testing the same thing”. This reflex has a name in the industry: benchmarketing, publishing a figure chosen for its marketing effect rather than for its methodological rigor.

Our response is not to refuse to measure — refusing to publish a figure indefinitely would be as dishonest as publishing an unsubstantiated one. It is to first document how we would measure, with what tools and under what conditions, before claiming to have measured anything. This is also what distinguishes a useful comparison (like our Aurabase vs Appwrite comparison, which documents verifiable architectural differences) from a comparison of performance figures without a common protocol.

This is particularly important for a technical lead or a CTO who must defend a backend choice in a technical committee: a figure that cannot be traced back to a method does not survive the first, somewhat insistent question. A documented protocol is self-defending — you can show the script, the tested version, and run the test again in front of someone if necessary.

#
Diagnosis

What Makes Most Backend Benchmarks Misleading

Two pitfalls come up systematically: comparing different topologies without reporting it, and measuring latency in a way that exactly hides the pauses that matter most to the user.

On the first point, PlanetScale explicitly documents its hardware parity constraint: each compared environment must run on computing resources (vCPU, RAM) equal to or greater than the reference instance, in the same cloud region (planetscale.com/benchmarks, “Telescope” methodology, accessed August 23, 2026). Without this discipline, a latency gap may simply reflect a larger machine — not a faster architecture.

The same principle applies to cache state and network topology. An instance that has just started (cold Postgres cache, empty connection pool, query plan not yet cached) responds structurally slower than an instance that has been running for an hour under stable load. A query from the same region as the database responds structurally faster than a cross-region query. Two benchmarks that specify neither are simply not comparable, even if they display identical units.

On the second point, the trap is called coordinated omission. HdrHistogram, the reference project on latency measurement created by Gil Tene, explains it this way: when a load generator waits for the response of a request before sending the next one (closed loop), a service pause automatically drops the number of requests sent during the pause — and therefore the number of high latency measurements recorded (github.com/HdrHistogram/HdrHistogram, accessed August 23, 2026). The project gives a concrete and quantified example: on a hypothetical system which samples its latency every 10 ms for 200 seconds, a single pause of 100 seconds in the middle of the test is enough to produce, without correction, a histogram where approximately 99.99% of the responses seem to fit under 1 ms - even though half of the real time has elapsed in this single pause.

Classic trap

A closed-loop load test that only sends a request after receiving the previous response systematically underrepresents long pauses. The p99 it displays may be better than the reality experienced by a real user — not because the system is fast, but because the measurement protocol "forgot" to send the queries while paused.

#
Statistics

Why the average lies: p50, p95, p99

An average latency may seem excellent when one in twenty requests takes five times as long. This is precisely what the percentiles reveal and what the average structurally conceals.

Mechanically, there is nothing mysterious about a percentile: sort all measured latencies in ascending order, then take the value at the corresponding position. Out of 1000 sorted queries, p50 is the 500th value, p95 the 950th, p99 the 990th. A single abnormally slow request among 1000 is enough to make the p99 move - it is precisely its sensitivity to rare cases that makes it useful, where this same isolated request has almost no effect on the average.

Telltale sign: the text report that pgbench — the official PostgreSQL benchmark tool — displays by default gives an average and a standard deviation, not percentiles (postgresql.org/docs/current/pgbench.html, accessed August 23, 2026). Its official documentation also warns: “Never believe any test that runs for only a few seconds” – never believe a test that only runs for a few seconds, which applies as much to the duration as to the chosen metric.

k6, the load tool that we use for level 2 of our suite, solves this with thresholds expressed in percentile: the p(95)<500 syntax defines a pass/fail criterion — 95% of requests must respond within 500 ms — directly in the test configuration (grafana.com/docs/k6, consulted on August 23, 2026).

p50 (median)Half of queries are faster than this valueCompletely hides the distribution tail
p951 in 20 queries is slowerArea where the first dissatisfied users appear
p991 in 100 queries is slowerMost sensitive to coordinated omission if the protocol is poorly designed
#
Checked in code

The 3 benchmark levels already present in our repository

Publishing a methodology without real tools would be just another form of theater. The benchmarks/ folder in the Aurabase repository already contains a 3-level suite, inspired in its structure by the public Supabase methodology - the tools exist, the measured and dated results do not yet exist.

3
TEST LEVELS
Micro, HTTP load, comparison
3
BENCHMARKED CRATES
aura-crypto, aura-db-adapters, aura-core
8
K6 SCRIPTS
7 wired to Makefile, 1 waiting

Level 1 — Micro-benchmarks Criterion.rs

Three crates of the Cargo workspace have dedicated CPU-bound benchmarks: aura-crypto (Argon2 hash, JWT HS256 — generation, validation and signature for PostgREST, AES-GCM encryption), aura-db-adapters (parsing filters and select in PostgREST format — eq., gte., in.(), relationship embeds), and aura-core (JSON serialization, schema_nameresolution, UUID validation).

libs/aura-crypto/benches/crypto_bench.rsrust
// Actual extract from the repository
let mut group = c.benchmark_group("jwt/hs256");

group.bench_function("generate", |b| {
    b.iter(|| jwt::generate_access_token(
        black_box(user_id), black_box(project_id),
        black_box("authenticated"), black_box(SECRET),
    ))
});

group.bench_function("validate", |b| {
    b.iter(|| jwt::validate_token(black_box(&token), black_box(SECRET)))
});

aura-db-adapters specifically measures the cost of parsing PostgREST format queries — four cases for filters (simple_4, complex_10, or_group, in_large_50 with 50 values) and four for select (single columns, *, one relation embed, five embeds). This is the kind of cost invisible in a global load test: a regression on the analysis of a complex or.(...) filter would change almost nothing in the p95 of a little-used endpoint, but would become measurable on an endpoint with high traffic - hence the interest in isolating it in a micro-benchmark rather than relying solely on level 2.

aura-core takes a different approach: rather than measuring raw time, it measures throughput (Throughput::Bytes) on the JSON serialization and deserialization of internal NatsRequest/NatsResponse messages exchanged between the gateway and the services — with three realistic payload sizes (a minimal request, a request with a JSON body nested, a 50-line list response).

libs/aura-core/benches/core_bench.rsrust
group.throughput(Throughput::Bytes(
    serde_json::to_vec(&small).unwrap().len() as u64
));
group.bench_function("NatsRequest/small", |b| {
    b.iter(|| serde_json::to_vec(black_box(&small)).unwrap())
});

Criterion.rs doesn't just time a loop. It first runs a warm-up phase to fill the CPU/OS caches, detects outliers with a modified version of Tukey's method (without excluding them from the dataset), calculates confidence intervals by bootstrapping on a large number of resampled samples, and detects performance regressions between two runs by Student's statistical test, with a configurable noise threshold — typically ±1% — to ignore variations that are not statistically significant (bheisler.github.io/criterion.rs/book/analysis.html, accessed August 23, 2026).

Locally reproducible

Each Criterion run generates a detailed HTML report in target/criterion/ — distributions, regression graphs, comparison to previous run. It is this relationship, not just a terminal line, that a serious methodology must make it possible to regenerate.

Level 2 — k6 load testing

Eight k6 scripts cover the gateway on the data plane side: health (latency baseline), auth-flow (register → login → refresh → logout), crud-read and crud-write, storage (upload/download), realtime-ws, breakpoint (load increase until failure) and supabase-compare. Seven are wired to a dedicated Makefile target — supabase-compare.js exists in the repository but does not yet have a target, a state of affairs that this article documents as is rather than disguising it.

benchmarks/k6/scenarios/crud-read.jsjavascript
export const options = {
  scenarios: {
    crud_read: {
      executor: 'ramping-vus',
      startVUs: 10,
      stages: [
        { duration: '30s', target: 50 },
        { duration: '1m', target: 200 },
        { duration: '30s', target: 0 },
      ],
    },
  },
  thresholds: THRESHOLDS_READ,
}

Shared configuration defines thresholds per operation type. These are pass/fail criteria that the test checks each time it runs — not results already measured:

Reading (GET)p95 < 500 ms · p99 < 1000 msConfig k6 (benchmarks/k6/lib/config.js)
Writing (POST/PATCH)p95 < 300 ms · p99 < 1000 msConfig k6
Auth (login/refresh)p95 < 300 ms · p99 < 1000 msConfig k6
Storage (upload/download)p95 < 500 ms · p99 < 2000 msConfig k6
Error rate, all scenarios< 1 %Config k6
An inconsistency found while writing this article

The README.md in the benchmarks/ folder documents a reading threshold of p95 at < 200ms ("Supabase SLO"), while the threshold actually applied in benchmarks/k6/lib/config.js — the one the test runs — is p(95)<500. The two files are derived from each other. This is a concrete example, found while reading the source code for this article, of why a protocol should have a single versioned source of truth rather than being documented in two places: without it, even a team that tries to be rigorous ends up publishing conflicting thresholds.

Level 3 — Direct PostgreSQL vs API comparison

A Python script (direct_vs_api.py) measures the actual overhead of the gateway + service layer by comparing direct psycopg2 requests to HTTP calls on the same operation — list, one-time read by id, filtered and sorted read. Each measurement follows a warm-up of 10 iterations before the timed loop, then calculates average, p50, p95, p99 and a throughput in operations per second.

benchmarks/comparison/direct_vs_api.pypython
def percentile(data, p):
    k = (len(data) - 1) * (p / 100)
    f = int(k)
    c = f + 1
    if c >= len(data):
        return data[f]
    return data[f] + (k - f) * (data[c] - data[f])

A second script (aurabase_vs_supabase.py) applies the same warmup and percentile calculation logic to a head-to-head comparison with a local Supabase instance (Supabase CLI, localhost:54321 by default) — same machine, same local network for both, exactly the environment parity discipline that PlanetScale documents for its own comparisons.

An orchestration script (collect_baseline.sh, target bench-baseline of Makefile) connects the three levels — Criterion on the 3 crates, a subset of the k6 scenarios (health and crud-read today, not all 8 yet), then the Python comparison — and writes logs, JSON and Criterion HTML reports to a folder unique timestamped: benchmarks/results/AAAAMMJJ_HHMMSS/. This is exactly the reflex of dated disclosure, in a single reproducible run, that the following section formalizes in a complete protocol.

#
Methodology

The protocol we will apply before publishing a figure

Eight commitments, each anchored in a practice already documented by a recognized third-party tool or project — not invented for the occasion.

  1. Preheating separate from measurement. Criterion.rs fills CPU/OS caches before timing; pgbench explicitly recommends never believing a run lasting just a few seconds.
  2. Fixed duration, not a fixed number of iterations. A load needs time to converge — this is the role of stages k6 and the -T flag of pgbench.
  3. Percentiles, never just the average — and active vigilance on coordinated omission if the load generator operates in a closed loop.
  4. Environment documented in detail: git commit of the service tested, version of PostgreSQL, hardware specification, version of the load tool. PlanetScale documents its exact TPCC parameters (TABLES=20, SCALE=250, ~500 GB) for precisely this reason — without these details, no one can reproduce a run.
  5. Timestamped and versioned results, never a single number engraved on a marketing page without a date. The current tooling is already written in a dated file; it will be necessary to extend this reflex to any publicly published measurement, with the hosting region documented like any other environmental variable (see our guide on EU hosting sovereignty, relevant as soon as a figure depends on a given region).
  6. Scripts and raw data published alongside the aggregate result, not just a final average. PlanetScale even invites readers to report a methodological error on a dedicated address – a posture that we find healthy and that we want to resume.
  7. Advertised throughput alongside latency, not just one or the other. A system can have excellent latency at low load and collapse in throughput as concurrency increases — that's exactly what our k6 suite's breakpoint scenario (scale-to-crash) is designed to reveal, and what Criterion's Throughput::Bytes microbenchmark measurement captures at the function level.
  8. Significant gap before announcing an improvement. A variation of a few percent between two runs may be measurement noise rather than a real gain — Criterion.rs calculates a probability that the observed difference is due to chance before qualifying it as regression or improvement. An isolated figure, without this verification, is just a statistical anecdote.
#
Editorial commitment

What we won't do

This list counts as much as the positive protocol above.

  • Comparing different topologies (self-hosted vs managed, cold vs pre-heated instance) without explicitly reporting it.
  • Retain the best run out of ten without mentioning the other nine.
  • Publish a figure without a date, without a service version, without a reproduction script.
  • Republish an existing marketing figure as long as it is not traced back to this protocol.
  • Comparing ourselves to a competitor on a raw performance figure if that competitor does not publish its own methodology in an equivalent manner — a figure versus silence is not a comparison, it's a slogan.
A concrete example, already corrected internally

A figure like “cold start less than 1 ms” was circulated without being backed by a reproducible benchmark. It is now treated internally as unsupported and should not be read as a measured characteristic of the product until any dated measurement, with published methodology, confirms it. This is precisely the kind of claim that this protocol exists to prevent from repeating.

#
Repeatable

The minimal protocol for benchmarking any backend

This protocol does not depend on any specific Aurabase tool — you can apply it to your own API today.

  1. Set the load before the tool: read-only, write, realistic mix for your application — not a generic ratio copied from another project.
  2. Explicitly separate the preheating phase from the measuring phase.
  3. Run the test long enough — minutes, not seconds.
  4. Measure in percentiles (p50/p95/p99), never on average alone.
  5. Verify that your load generator is not in closed loop, or correct the coordination omission in the analysis.
  6. Isolate the environment under test — no noisy neighbors, no competing background tasks.
  7. Publish the tested version, date, hardware spec, and script — not just the final result.

On a bare Postgres base, this protocol takes one command pgbench — 20 concurrent clients distributed over 4 threads, for 5 minutes, with a progress report every 10 seconds:

terminalbash
# Initialize the test dataset (scale factor >= number of clients)
pgbench -i -s 50 ma_base

# -c concurrent clients, -j threads, -T duration in seconds, -P reporting interval
pgbench -c 20 -j 4 -T 300 -P 10 ma_base
#
Tools

Reference tools, by level

Five tools, each suited to a different level of the stack—none replaces the others.

Microphone (function)Criterion.rsPure CPU, bootstrap statistics
SQL QuerypgbenchTPC-B-like transaction, tps and latency
HTTP/WS loadk6 (Grafana)Percentiles, pass/fail thresholds
OLTP at scalesysbench + TPCC (Telescope methodology)QPS, cost per performance
Measurement correctionHdrHistogramCompensates for coordinated omission
#
Frequently Asked Questions

FAQs

Why isn’t Aurabase publishing benchmark numbers yet?+
Because no performance figures have been measured according to a published and reproducible protocol to date. A figure like “cold start less than 1 ms”, for example, was circulated without being backed by a reproducible benchmark: it is today treated as unsubstantiated and should not be read as a measured characteristic of the product. This article documents the protocol we will follow before publishing a result, precisely to avoid repeating this kind of claim.
What is a p95 or p99 percentile, and why not the average?+
The p95 is the response time below which 95% of requests are found — 1 in 20 requests is therefore slower. The p99 pushes this threshold to 1 request in 100. The average hides these slow requests because it dilutes them in the mass of fast requests; percentiles isolate the tail of the distribution that users actually notice.
What is “coordinated omission”?+
This is a measurement bias described by the HdrHistogram project: when a load tool waits for the response to a request before sending the next one (closed loop), a service pause mechanically reduces the number of slow requests recorded during this pause. The end result can show much better latency than what a real user experienced.
Can we reproduce these tests ourselves?+
The protocol described here — percentiles, separate preheating, documented environment, dated results — is applicable to any API, with public tools (k6, pgbench, Criterion.rs, sysbench). The internal Aurabase tooling (the benchmarks/repository folder) is currently used for development and is not yet packaged as a one-click public suite. Create a project to test your own load on the Aurabase API with your own k6 scripts.
What is the difference between a measured percentile and an SLA threshold?+
A percentile (p95, p99) is a statistic calculated after the fact on real measurements. An SLA threshold (or a k6 threshold like p(95)<500) is a target set in advance, which the test verifies in pass/fail mode. Confusing the two leads to presenting an unachieved objective as an obtained result – this is precisely the distinction that this protocol requires to be kept explicit with each published figure.
Why measure throughput in addition to latency?+
A system can respond quickly at low load and see its latency suddenly degrade once a concurrency threshold is exceeded — latency alone does not show where this threshold is. Measuring throughput (requests or bytes per second) alongside latency reveals this tipping point, which is something our k6 suite's breakpoint scenario is specifically designed to find.

READY TO DEPLOY?

Your backend in five minutes.

No credit card required · 500 MB free · 50,000 MAU