This article relies solely on third-party, dated sources — never on an invented Aurabase cipher. No p99 latency specific to our backend is included. We write our application core in Rust, without garbage collector: a fact that can be verified directly in the code, workspace Cargo and axum services. However, we have not yet published a reproducible p99 benchmark methodology to demonstrate this in figures. This text explains a mechanism, not a measured result.
The essentials
- p99 measures the slowest query out of a hundred — the place where a garbage collector (GC) pause hurts the most, not on average (Dean & Barroso, “The Tail at Scale,” Google, 2013).
- A GC interrupts the entire program (“stop-the-world”) to free unused memory. Rust has no GC: memory is freed at the precise moment when a value goes out of scope, verified by the compiler.
- Discord documented a cache service in 2020 where Go triggered a garbage collection cycle at least every two minutes, with each cycle causing a spike in latency (Discord Engineering Blog).
- Reducing a GC pause takes years of engineering, even at Google: the GB collector went from 300-400 ms to 500 µs between 2015 and 2018, without ever reaching zero (go.dev).
- The backend core of Aurabase is written in Rust, without garbage collector — verified in the code. No p99 Aurabase latency figures have been published to date: this remains a mechanism, not a measurement.
Why p99 is not average
An average hides the essential. If 99 requests out of 100 respond in 5 ms and only one takes 500 ms, the average remains low. But one user in a hundred experiences a wait a hundred times longer. The p99 measures exactly this request: the slowest hundredth, the one that violates your SLA while your average latency dashboard remains green.
At Google, Jeffrey Dean and Luiz André Barroso formalized this problem in “The Tail at Scale” (Communications of the ACM, vol. 56, 2013). Their observation, often cited since: “Temporary high latency episodes which are unimportant in moderate size systems may come to dominate overall service performance at large scale”. In short: occasional episodes of latency, negligible on a small scale, end up dominating the perceived performance of a distributed system.
A backend that processes thousands of requests per second is bound to send, at one point or another, a request that falls during a GC pause. On a large scale, this is not an uncommon case. This is a statistical certainty.
What a garbage collector does, and why it pauses everything
A garbage collector (GC) continuously tracks the living objects of a program — those still referenced somewhere — and frees the memory of objects that have become inaccessible. This tracking is called tracing: the GC goes through the reference graph, marks what is still used, then sweeps the rest.
The problem: traversing this graph while the program continues to create new references produces inconsistent results. The historical answer, still used as a last resort by many modern GCs, is stop-the-world — the entire program pauses while marking and scanning. The larger the heap, the longer the pause tends to be: its duration depends on the size of the live data, not the current workload.
Most modern GCs use a generational strategy: they assume that the majority of objects die young. Recent allocations are therefore scanned often, but quickly, in a small memory area. Objects that survive several cycles migrate to a larger area, scanned infrequently — but when that area needs to be cleaned, the associated pause grows with its size. It is this “major” break, not the small “minor” breaks, that dominates the p99 of a high-traffic, high-allocation service.
Modern concurrent and generational GCs reduce the frequency and duration of these breaks by working in parallel with the program. But almost all of them keep a stop-the-world fallback mechanism for borderline cases — and reducing it takes years of engineering. Section 04 gives a quantified and sourced example.
Discord, 2020: a GC break becomes a production incident
In February 2020, engineer Jesse Howarth published a post that has become a reference in the industry: “Why Discord is switching from Go to Rust” (Discord Engineering Blog). The relevant service, Read States, manages message read status for millions of users — tens of millions of entries per cache — with hundreds of thousands of updates per second.
The diagnosis is direct, quoted as is in the article: “Go will force a garbage collection run every 2 minutes at minimum”. In other words, Go triggers a garbage collection cycle at least every two minutes on this service — and each cycle produces a spike in latency visible in the team's graphs.
The team first reduced the cache size to smooth out the spikes. The compromise remained unfavorable: fewer GC pauses, but more cache miss requests falling on the database — therefore a degraded overall p99 elsewhere. The basic fix was the rewriting of the service in Rust, without a garbage collector to monitor.
The post sparked a lively technical debate: more than 1,580 points and 642 comments on Hacker News on the same day of its publication (February 4, 2020) — a sign that the problem goes far beyond the Discord case.
Three years of engineering at Google to drop a pause from 400 ms to 500 µs
The Go garbage collector illustrates the scale of the effort required to tame a GC pause — even with the resources of a dedicated team at Google. Rick Hudson, Go GC technical lead, documented this story in two official Go blog posts.
| Before August 2015 | 300-400ms | Historical Go collector, before redesign |
|---|---|---|
| August 2015 · Go 1.5 | 30-40ms | First competing collector, target < 10 ms set |
| 2016 · Go 1.6 | < 10 ms (SLO held) | Initial objective achieved in production |
| March 2017 · Go 1.8 | under the millisecond | Removing stop-the-world stack scan |
| August 2017 · Go 1.9 | 100-200 µs (mark) | New informal benchmark mentioned by the team |
| 2018 · SLO announced | 500 µs per cycle | Service objective formalized by Rick Hudson |
Source: “Getting to Go: The Journey of Go’s Garbage Collector”, go.dev, July 12, 2018; and “Go GC: Prioritizing low latency and simplicity”, go.dev, August 31, 2015.
Three years of dedicated work has dropped the typical break by a factor of a thousand. But the pause never disappeared: it's a service objective (SLO), not an absolute guarantee of zero. A GC tracing must, by construction, traverse a graph of living objects from time to time. The only adjustable variable is the frequency and duration of this journey — not its existence.
This choice of priority is not neutral. Go primarily targets network services and web backends, where a pause of several hundred milliseconds directly breaks the user experience — hence the massive effort invested in latency rather than raw GC throughput. Other managed runtimes inherited different trade-offs, shaped by their historical use cases, before making up ground with their own low-pause collectors. The common point remains the same: they all start from GC tracing, therefore from a pause mechanism to be minimized – never to be eliminated by construction.
Why Rust doesn't have this problem by construction
Rust doesn't reduce GC pauses: it eliminates the mechanism that causes them. The compiler tracks, upon compilation, who owns each memory value — this isownership. When the owner of a value goes out of scope, Rust automatically inserts the call that frees that memory, in the same place in the binary code. This mechanism is called RAII (Resource Acquisition Is Initialization): the release is deterministic, not scheduled by a garbage collector running in the background.
Alexandru Nedelcu, author of a technical blog recognized in the Scala/Rust ecosystem, summarizes the trade-off in a recent article: “The trade-off that Rust makes is one of ease-of-use, in preference for performance with predictable latency and safety” (alexn.org, July 21, 2026). Rust trades some of the simplicity of writing for predictable latency.
The same article summarizes why modern GCs are not always enough: “Modern GCs try to do their work incrementally and concurrently, without affecting the program. But their ability is limited, falling back to a stop-the-world GC cycle that freezes the whole program, thus affecting latency”.
Here is the mechanism in around ten lines — a generic example, not an extract from the Aurabase code:
Important nuance: not everything is free. Reference-counting types (Rc, Arc) add a small cost to each clone and release. This cost remains local and deterministic. There is never a pause that freezes the entire program while it goes through a memory heap.
Useful clarification for an asynchronous backend: the Rust async runtime (tokio, used by all Aurabase services) has nothing to do with a garbage collector. It schedules cooperative tasks on a pool of threads, but never iterates through a live object graph to free memory. Confusion is common coming from ecosystems where the asynchronous runtime and the GC are managed by the same virtual machine.
What this changes for a high traffic backend
On a backend that serves thousands of concurrent requests, the absence of GC removes a variable from the equation p99. No more need to size a memory heap, adjust the generations of a collector, or monitor a cycle that can fall at the worst time. The latency of an individual request depends on its own work, not on some unpredictable global event elsewhere in the program.
The backend core of Aurabase applies this principle: all services (aura-gateway, aura-auth, aura-db, aura-realtime, aura-storage, aura-functions, aura-ai…) are written in Rust, organized in a single Cargo workspace. This can be verified directly in the repository:
Actual extract from Cargo.toml, workspace edition 2021, resolver v2 — verified in the Aurabase repository.
What this fact does not prove, at this stage: a p99 latency figure measured for Aurabase. We have not yet published a reproducible benchmark methodology for our own backend — this is a work in progress, not a result available today. The absence of a garbage collector is a mechanism verified in the code. This alone is not evidence of measured p99 latency. Keep this distinction in mind in the face of any marketing argument on the subject, including ours — see our technical comparison Aurabase vs Supabase for details of the architecture.
Measuring a p99 correctly requires its own discipline: representative load conditions, percentiles calculated over a sufficiently wide sliding window, and a test environment close to production. Publishing a figure without this methodology is like publishing a marketing figure. This is exactly what we refuse to do in this article.
What the absence of GC does not solve
Removing the garbage collector eliminates just one source of tail latency — not all. A Rust backend can still show a degraded p99 due to network waiting, a saturated Postgres connection pool, a contested database lock, a poorly indexed SQL query, or a slow third-party API call. The mechanism described in this article removes a structural cause. It does not provide immunity against others.
At Aurabase for example, each service communicates with Postgres via a connection pool (sqlx) and with other services via NATS JetStream. An undersized pool, a slow-to-consume NATS subscription, or an SQL query without a suitable index each produce their own latency spike — regardless of the absence of a garbage collector.
The practical conclusion: the absence of GC is a good architectural reason to choose a Rust backend for a p99-aware system. This is not, in itself, a guarantee of latency – neither at Aurabase, nor elsewhere. The method that matters remains the same: measure, publish the methodology, then correct what the measurements reveal. If you're migrating from a backend with GC, our Supabase to Aurabase migration guide details what's changing and what's staying the same.