PRODSovereign European BaaS platformOpen Dashboard →

Performance · 10 min read

Gateway rate limiter and circuit breaker latency

Affane Daylami · Fondateur · March 11, 2026

Back to blog

A poorly placed rate limiter can add tens of milliseconds to each request. Well designed, it adds less than one. Aurabase's Rust gateway (aura-gateway) implements both mechanisms — distributed rate limiting and circuit breaker — at the heart of its middleware stack. Here's what the technical literature says, and what the Aurabase code does precisely.

This English text was generated automatically from the French original and has not been reviewed yet.

The essentials

A properly implemented distributed rate limiter (atomic counting, local fallback cache) typically adds less than a millisecond per request according to several specialized sources. The Aurabase gateway follows exactly this pattern: the governor crate for local anti-DoS, an atomic Lua Redis script for distributed counting between instances, a fallback moka cache if Redis is unavailable. The breaker circuit reduces the perceived latency in the event of a failure rather than adding it — it short-circuits the wait for a complete timeout. No Aurabase latency figures have been published to date: here's why, and how the mechanism actually works.

#
Why these two middlewares

Anti-DoS and anti-cascading failure, two distinct problems

Rate limiting protects your backend from excessive traffic, legitimate or not — it answers the question “does this caller have the right to send this request now?” ". The circuit breaker protects your backend from an already down downstream service — it responds to “has this service ever shown to become unresponsive, should we even try?” ". Confusing them leads to undersizing one or the other.

On the Aurabase gateway, both live in the same middleware stack but on different floors: IP rate limiting runs before authentication (pure anti-DoS, no billing SQL queries issued for unauthenticated traffic), while the circuit breaker protects outgoing calls to internal services or LLM providers.

#
What the published benchmarks say

Cost depends entirely on implementation, not principle

According to Tyk and the APISIX ecosystem guides, well-designed distributed counting on Redis (atomic operations, Lua script, native TTL) typically adds 1-3 ms of latency, with sub-millisecond p99 impact in the best-optimized cases. Conversely, Zuplo documents that a poorly placed centralized rate limiter can add tens of milliseconds to each request — this discrepancy is reflected directly in your p99.

Third-party figures, different methodologies

These ranges come from public technical guides (Tyk, Zuplo, Apache APISIX ecosystem), not from a common measurement protocol. They indicate an order of magnitude and an architectural principle — atomic and local counting rather than synchronous and centralized — not a figure to be reproduced as is on a different infrastructure.

#
Verified implementation

How rate limiting actually works in aura-gateway

The Aurabase gateway uses the governor crate (token bucket algorithm) as a local fallback limiter, coupled with distributed counting via Redis to share the state between several instances of the gateway. Distributed counting goes through a Lua script executed atomically on the Redis side (INCRBY + EXPIRE in a single network operation), not through a read-then-write round trip which would introduce a concurrency window.

If Redis becomes unavailable, the gateway automatically switches to the local limiter governor, with a cache moka (max 1,000,000 entries, expiration at 300 seconds of inactivity) to avoid recreating a limiter on each request. This fallback is observable: each toggle increments a dedicated Prometheus counter, so that the team knows when the effective limit becomes N instances × limit again (silent inter-instance isolation degradation otherwise).

Two layers of rate limiting, not just one

The gateway distinguishes anti-DoS rate limiting by IP (before authentication) from rate limiting by project/API key/user quota (after authentication, with JWT claims available) — two separate middlewares in the stack, each with its own granularity.

#
Verified implementation

The circuit breaker: three states, a sliding window

The Aurabase breaker circuit (aura_core::circuit_breaker, shared between the gateway and aura-ai for fallback between LLM providers) follows three classic states — Closed, Open, HalfOpen — with a count of failures over a sliding time window rather than a counter that never resets.

An implementation detail that matters in production: switching to HalfOpen only allows one probe request at a time (anti thundering-herd) — without this guard, all pending requests would rush simultaneously to the service that has just reopened, immediately recreating the failure we were trying to avoid.

#
Order in stack

Where these middlewares run in the gateway

The Aurabase gateway middleware stack follows a precise order, verified in the code: request identifier → access log → rate limiting by IP → authentication → rate limiting by actor/quota → circuit breaker → proxy to the target service. Each stage rejects unwanted traffic as early as possible, before incurring a higher processing cost downstream.

Read the full article: data plane vs management plane in the Aurabase gateway

#
Methodology

Why no Aurabase figures are published here

The implementation is verified line by line in the gateway source code. But no reproducible measurement protocol has yet been run and published on this specific infrastructure — publishing an unmeasured figure would be repeating the mistake of unsourced marketing numbers that we refuse to reproduce. See our full benchmark methodology for what we require before releasing a performance figure.

#
Frequently Asked Questions

FAQs

Does rate limiting still add noticeable latency?+
This is entirely implementation dependent. A well-designed distributed counter (atomic Lua script on Redis) typically adds less than a millisecond to p99 according to benchmarks published by Tyk and the APISIX ecosystem. A poorly placed rate limiter in the stack can, on the contrary, add tens of milliseconds.
Has the Aurabase gateway published a latency figure for its rate limiting?+
No. The implementation (crate governor for local fallback, Lua Redis script for distributed counting, mocha cache) is verified in the code, but no reproducible latency benchmark has been published to date on this specific infrastructure.
Why does a breaker circuit reduce perceived latency rather than add it?+
Because an open circuit breaker short-circuits the call to a down service instead of waiting for its complete timeout. An immediate failure response costs less in perceived latency than a request that waits several seconds before failing.

READY TO DEPLOY?

Your backend in five minutes.

No credit card required · 500 MB free · 50,000 MAU