The essentials
A properly implemented distributed rate limiter (atomic counting, local fallback cache) typically adds less than a millisecond per request according to several specialized sources. The Aurabase gateway follows exactly this pattern: the governor crate for local anti-DoS, an atomic Lua Redis script for distributed counting between instances, a fallback moka cache if Redis is unavailable. The breaker circuit reduces the perceived latency in the event of a failure rather than adding it — it short-circuits the wait for a complete timeout. No Aurabase latency figures have been published to date: here's why, and how the mechanism actually works.
Anti-DoS and anti-cascading failure, two distinct problems
Rate limiting protects your backend from excessive traffic, legitimate or not — it answers the question “does this caller have the right to send this request now?” ". The circuit breaker protects your backend from an already down downstream service — it responds to “has this service ever shown to become unresponsive, should we even try?” ". Confusing them leads to undersizing one or the other.
On the Aurabase gateway, both live in the same middleware stack but on different floors: IP rate limiting runs before authentication (pure anti-DoS, no billing SQL queries issued for unauthenticated traffic), while the circuit breaker protects outgoing calls to internal services or LLM providers.
Cost depends entirely on implementation, not principle
According to Tyk and the APISIX ecosystem guides, well-designed distributed counting on Redis (atomic operations, Lua script, native TTL) typically adds 1-3 ms of latency, with sub-millisecond p99 impact in the best-optimized cases. Conversely, Zuplo documents that a poorly placed centralized rate limiter can add tens of milliseconds to each request — this discrepancy is reflected directly in your p99.
These ranges come from public technical guides (Tyk, Zuplo, Apache APISIX ecosystem), not from a common measurement protocol. They indicate an order of magnitude and an architectural principle — atomic and local counting rather than synchronous and centralized — not a figure to be reproduced as is on a different infrastructure.
How rate limiting actually works in aura-gateway
The Aurabase gateway uses the governor crate (token bucket algorithm) as a local fallback limiter, coupled with distributed counting via Redis to share the state between several instances of the gateway. Distributed counting goes through a Lua script executed atomically on the Redis side (INCRBY + EXPIRE in a single network operation), not through a read-then-write round trip which would introduce a concurrency window.
If Redis becomes unavailable, the gateway automatically switches to the local limiter governor, with a cache moka (max 1,000,000 entries, expiration at 300 seconds of inactivity) to avoid recreating a limiter on each request. This fallback is observable: each toggle increments a dedicated Prometheus counter, so that the team knows when the effective limit becomes N instances × limit again (silent inter-instance isolation degradation otherwise).
The gateway distinguishes anti-DoS rate limiting by IP (before authentication) from rate limiting by project/API key/user quota (after authentication, with JWT claims available) — two separate middlewares in the stack, each with its own granularity.
The circuit breaker: three states, a sliding window
The Aurabase breaker circuit (aura_core::circuit_breaker, shared between the gateway and aura-ai for fallback between LLM providers) follows three classic states — Closed, Open, HalfOpen — with a count of failures over a sliding time window rather than a counter that never resets.
An implementation detail that matters in production: switching to HalfOpen only allows one probe request at a time (anti thundering-herd) — without this guard, all pending requests would rush simultaneously to the service that has just reopened, immediately recreating the failure we were trying to avoid.
Where these middlewares run in the gateway
The Aurabase gateway middleware stack follows a precise order, verified in the code: request identifier → access log → rate limiting by IP → authentication → rate limiting by actor/quota → circuit breaker → proxy to the target service. Each stage rejects unwanted traffic as early as possible, before incurring a higher processing cost downstream.
Read the full article: data plane vs management plane in the Aurabase gateway
Why no Aurabase figures are published here
The implementation is verified line by line in the gateway source code. But no reproducible measurement protocol has yet been run and published on this specific infrastructure — publishing an unmeasured figure would be repeating the mistake of unsourced marketing numbers that we refuse to reproduce. See our full benchmark methodology for what we require before releasing a performance figure.