PRODSovereign European BaaS platformOpen Dashboard →

Engineering · 10 min read

Designing a dual-plane API gateway in Rust

Affane Daylami · Fondateur · August 6, 2026

Back to blog

A gateway that serves both your users' SDK traffic and your back office admin traffic poorly protects both at the same time. The Aurabase gateway (aura-gateway, Rust/axum) resolves this tension upstream: two distinct Router, two ports, two authentication models, a single shared state. This guide describes this data-plane/management-plane pattern as it actually exists in the code — routes, middleware order, rate limiting, circuit breaker and proxy — not an idealized version of an architectural diagram.

This English text was generated automatically from the French original and has not been reviewed yet.

The essentials

Two Router axes on two ports (8080 data plane, 8090 management plane), built from the same shared AppState. The data plane requires an API key on any route; Plane management requires a dedicated JWT audience console — the two mechanisms never overlap. Rate limiting applies twice: by IP before authentication, then by authenticated actor after. The circuit breaker is not a global middleware: it is an object per service (and per dedicated target for PostgREST), invoked directly in the proxy code. And the final proxy changes transport depending on the route — sometimes according to the HTTP method or a database lookup: NATS request/reply for most of the traffic, direct HTTP flow for storage, the three realtime variants and the dedicated Postgres CRUD.

#
The problem

A single gateway, two very different audiences

Data plane traffic comes from the SDK or a client app: a volume of anonymous requests or requests authenticated by API key, with an abuse profile close to any public API. The traffic management plane comes from the Studio — the administration interface of a project — and carries sensitive operations: project creation, key rotation, reading of a tenant's logs. The two share a common target (proxiing to the same internal services: aura-auth, aura-db, aura-storage, etc.) but not the same risk surface.

Passing both through the same router requires choosing between two bad options: either the Studio CORS inherits the wildcard necessary for the public SDK (Access-Control-Allow-Origin: *), or the SDK inherits a restricted list of origins designed for an internal dashboard. The aura-gateway code sorts out this tension from main.rs: two separate Router, each with its own CorsLayer — wildcard authorized on the data plane side, refused and logged as an error on the management plane side.

#
Step 1

Two axum Routers, one shared AppState

The separation is not a separate deployment: the two plans run in the same process, on the same AppState (Postgres pools, NATS client, Moka caches, circuit breakers). Only the construction of Router differs, via two dedicated functions each called once at startup and served by two separate TcpListener.

gateway/server.rsrust
// Two ports, two Routers, one AppState

let data_addr: SocketAddr = format!("{}:{}", config.host, config.port).parse()?;
let mgmt_addr: SocketAddr = format!("{}:{}", config.host, config.management_port).parse()?;

let data_listener = TcpListener::bind(data_addr).await?;
let mgmt_listener = TcpListener::bind(mgmt_addr).await?;

// Same state, two distinct route graphs
let data_app = build_data_plane_router(state.clone(), data_cors);
let mgmt_app = build_management_plane_router(state);

tokio::join!(
    axum::serve(data_listener, data_app),
    axum::serve(mgmt_listener, mgmt_app),
);

The two route graphs start from the same base, service_routes(): the same proxy handlers (db_proxy, storage_proxy, functions_proxy…) are mounted on both plans, with additional routes specific to each. Reusing the same handlers avoids a double implementation of the proxy; diverging only on middleware avoids duplicating business logic to gain a security boundary. If your backend is itself structured as a multi-service Cargo workspace, see our Cargo workspace architecture guide — the gateway is just one crate among others in this division.

#
Step 2

Authentication diverges upon entry

On the data plane, the API key is mandatory on any route, except a handful of truly public paths (/health, JWKS, registration endpoints). It travels as a apikey or X-API-Key header — or, for WebSocket and SSE streaming routes only, as a ?apikey=parameter. The code explicitly prohibits this last mode for a service_role key: a URL key leaks in access logs, OTel traces and the Referer header. A JWT remains optional on the data plane side: without it, the caller remains anon; with it, it becomes authenticated.

On the management plane, the API key does not exist: only a JWT console is accepted, whose audience must be exactly aurabase-control. The role is not carried by the token itself — it is recalculated on each request from the user's membership in the organization that owns the project, inherited via the project → organization relationship.

Token requiredAPI key (apikey / X-API-Key), alwaysJWT console (Authorization: Bearer), always
Role elevationOptional JWT: anon → authenticatedRBAC inherited from the organization (owner/admin/developer/viewer)
Key in query stringTolerated on WS/SSE only, never for service_roleNot applicable
Expected audienceThe targeted project (UUID of the path)fixed "aurabase-control"
CORSWildcard * allowedWildcard refused, Studio origins only
#
Step 3

The Real Order of Middleware (And Why It Matters)

axum stacks middleware with successive .layer() calls — and the rule that governs the order of execution is surprising in practice: the LAST .layer() placed becomes the OUTTERmost layer, therefore the first traversed by an incoming request, and the last to see the response leave. A linear reading of the file therefore gives the reverse order of the actual execution order.

gateway/router.rsrust
// Written as follows (actual extract, file order):

service_routes(...).merge(data_plane_extra)
    .layer(metrics_auth_middleware)      // (1) placed 1st → the innermost
    .layer(rate_limit_actor_middleware)  // (2)
    .layer(data_plane_auth_middleware)   // (3)
    .layer(rate_limit_middleware)        // (4)
    .layer(request_id_middleware)        // (5)
    .layer(AccessLogLayer)               // (6)
    .layer(prometheus_layer)             // (7)
    .layer(TraceLayer)                   // (8)
    .layer(RequestBodyLimitLayer)        // (9)
    .layer(security_headers_middleware)  // (10)
    .layer(cors)                         // (11) placed last → outermost

// An incoming request therefore passes through (11) → (1), never (1) → (11).
A concrete effect of this order

The request_id middleware only sets the X-Request-Id header on the RESPONSE, never on the incoming request. As AccessLogLayer is placed after it in the file - therefore more external, therefore traversed before - its capture of the request_id field reads the header as the client sent it, not the identifier generated further in the chain. If the caller has not provided any X-Request-Id, the access-log line leaves with an empty field, while the response returned carries a freshly generated UUID. Not a hidden defect — a reminder that the order in which a .layer() string is written does not guarantee anything about the logical order that we attribute to it.

#
Step 4

Rate limiting: IP first, actor then

Rate limiting is applied in two distinct passes, at two different times in the chain. The first runs BEFORE authentication and limits by IP address — a generic anti-flood filter, active even on public roads: without it, an unauthenticated flow can hammer an expensive endpoint, such as a log aggregation, without ever triggering a JWT check. The second runs AFTER authentication and limits by actor — API key or user — using the claims that the authentication has just injected: this is the real product quota, the one that counts for billing and plans.

The implementation relies on the governor crate (token bucket) for local calculation, with a sliding window Lua Redis script for distribution between gateway instances, and a local fallback (Moka cache) if Redis is unavailable. Repository defaults: 100 requests/second, burst of 1000.

#
Step 5

The circuit breaker is not a layer, it is an object per target

Unlike the rest of the chain, the circuit breaker does not appear in ANY .layer(). AppState carries one CircuitBreaker instance per service (auth, db, realtime, storage, functions, notifications, ai, provisioner, control), and it is the proxy code itself — not the router — that calls try_acquire_probe() before attempting the request, then record_success() or record_failure() depending on the outcome.

The PostgREST case is distinct: projects in dedicated topology (project-specific Postgres and PostgREST) do not have a shared fault domain — each PostgREST process is its own target. The gateway therefore maintains a table of circuit breakers indexed by resolved target, populated on the fly and purged every 60 seconds by a sweep which removes inactive entries: without this purge, each new dedicated project would add an entry which never disappears.

gateway/circuit_breaker.rsrust
let Some(probe) = circuit_breaker.try_acquire_probe() else {
    return Err(ServiceUnavailable);
};

// …NATS query attempt(s), with bounded retry…

match resultat {
    Ok(Ok(_))     => match probe.take() { Some(p) => p.record_success(), _ => {} },
    Ok(Err(_))    => match probe.take() { Some(p) => p.record_failure(), _ => {} },
    Err(_timeout) => {} // probe not consumed → returned by Drop
}

The probe token is returned as Drop if it is never explicitly consumed — useful when all attempts at a request time out without reaching a branch that would have freed it. And the automatic replay is only triggered on a strict proof of non-delivery on the NATS side (NoResponders): a simple gateway timeout does not prove anything about the real delivery of the request, and replaying it could execute it twice.

The final proxy does not speak a single protocol backwards, and the choice is not fixed by route: it can depend on the HTTP method, or even on a database lookup. For most of the traffic (auth, functions, notifications, control, and the majority of db), the gateway serializes the HTTP request in a NATS envelope and sends it as request/reply to a subject dedicated to the service — a round trip without TCP handshake, documented in the code as significantly faster than a classic HTTP proxy for this RPC type traffic.

Storage, the three variants of realtime (WebSocket, SSE, and REST for broadcast/channels/presence) and — conditionally — Postgres CRUD requests exit this path and go through a live, pooled HTTP client. Storage made this choice explicitly: encoding a binary body in a NATS envelope requires serializing it, loading it entirely into memory on both ends, and staying under the NATS message size cap — a real cost for large objects. WebSocket and SSE simply do not tolerate request/reply semantics: a protocol upgrade and a flow that remains open have no NATS equivalent.

The most interesting case is /v1/db/*, whose handler decides itself on each request: the management routes (schema, policies, raw SQL) always go in NATS to aura-db, an PUT always goes in NATS (PostgREST returns 405 on a complete replacement), a MongoDB project always goes in NATS — and only a CRUD on a Postgres project with a dedicated PostgREST instance resolved goes in Direct HTTP. If this dedicated instance is not resolved, the gateway responds 503 rather than falling back on a shared PostgREST: fail-closed assumed, not a degraded silent fallback. The security headers (strict CSP, no CORS credentials) apply uniformly to all these paths, placed at the very end of the chain, before the response leaves the gateway.

#
Step 7

A timeout budget per route, not a global timeout

The gateway applies a TimeoutLayer PER GROUP of routes rather than a global timeout — a choice linked to the same stacking mechanics as the middleware order. The Edge functions route needs a much longer budget than the rest (a function can legitimately run for several minutes): the repository default is 30 seconds for the majority of routes, compared to 380 seconds for /v1/functions/*.

Stacking a single global TimeoutLayer on top of everything would have cut both groups at the same limit: it is always the shortest timeout placed in the outermost position that wins, regardless of a longer timeout placed further inside. The only way to grant functions a separate budget is therefore that they NEVER enter into a common wrap: each branch of routes carries its own TimeoutLayer, placed before the merger of the two routers — and no global timeout is applied afterwards.

#
To remember

Reproduce this pattern elsewhere: the checklist

  1. Separate by PLAN (exposure surface), not by service: a compromised public SDK should never reach the CORS origin list of your admin dashboard.
  2. Keep a single shared state rather than two separate deployments — duplicating business logic costs more than duplicating a router.
  3. Check the ACTUAL middleware order by tracing it from the last .layer(), never from the linear reading of the file.
  4. Separate rate limiting per IP (before auth) from quota per actor (after) — otherwise an unauthenticated flow forces costly verification without limit.
  5. Place the breaker as close as possible to the actual network call, in the proxy — and size it by target when the fault domain is not shared.
  6. Only replay a request on proof of non-delivery, never on a simple timeout.
  7. Give each route group its own timeout budget set before merging routers — never a global TimeoutLayer that would overwrite the longest budget.

READY TO DEPLOY?

Your backend in five minutes.

No credit card required · 500 MB free · 50,000 MAU