Our article on PostgREST compatibility details what the server covers functionally (filters, embedding, RPC, RLS) and what it leaves up to you. This one comes from somewhere else. It documents, with Aurabase code and the official PostgREST documentation as sources, where and why PostgREST actually plateaus at scale. We are not reproducing here a load bank that we have not run ourselves. Our benchmark methodology explains why an isolated figure, without a published protocol, does not seem reliable to us.
The essentials
- PostgREST itself is lightweight: 50 to 250 millicores of CPU, 64 to 128 MB of RAM per replica on dedicated Aurabase instances. Raw HTTP throughput is almost never the limiting factor in production.
- The real ceiling is the Postgres connection budget:
PGRST_DB_POOL× replicas. Verified in the Aurabase code: 20 connections per project on the dedicated level (10×2), 4 on the shared level (2×2). This is a deliberate choice to fit more tenants on the samemax_connections. Prefer: count=exactforces expensive MVCC scanning on large tables. PostgREST documents two cheaper alternatives:count=plannedandcount=estimated, costing an approximate total.- A
db-max-rowsceiling (1000 lines by default at Aurabase) truncates a response WITHOUT reporting it inContent-Range(measured in real conditions, detailed below). - After a DDL migration, the PostgREST schema cache reloads asynchronously. The Aurabase gateway retries up to 8 times (around 3.5 seconds cumulative in the worst case) before giving up, a behavior documented directly in the code.
What a PostgREST benchmark measures, and what it doesn't measure
An HTTP throughput test on PostgREST mainly measures Postgres, rarely PostgREST. The server is a thin layer of translation in front of the base. In the vast majority of real-world loads, response time is dominated by the SQL query executed, not the process that generated it.
The PostgREST project maintains a dedicated repository for this topic, PostgREST/postgrest-benchmark on GitHub, which tracks throughput variations from release to release rather than publishing an isolated marketing figure. We have neither performed nor republished it here. Its results depend on hardware, schematic size, and scenario tested, exactly the variables that our own benchmark protocol requires to be documented before quoting a figure.
Below PostgREST, it's pgbench which measures the layer that really matters: SQL transaction time under concurrent load. This is the official PostgreSQL benchmark tool (postgresql.org/docs/current/pgbench.html, accessed August 24, 2026). Rather than reproducing this protocol here, this article documents four concrete architectural limitations of PostgREST in production, each verified in the Aurabase source code or in the official project documentation.
The actual footprint of a PostgREST instance at Aurabase
Each Aurabase Postgres engine project receives two dedicated PostgREST replicas, co-located with its cluster. The Kubernetes manifest that deploys them sets modest resources.
What these replicas really consume is not CPU: they are connections to the Postgres primary. Each PostgREST instance connects directly to the primary (-rw), without going through the PgBouncer pooler deployed for the tenant. This choice is already detailed in our article on PostgREST compatibility: the LISTEN/NOTIFY schema reloading mechanism requires a persistent connection, incompatible with a pooler in transaction mode. What this article adds: how much does it actually cost, in terms of connections, and where it peaks.
The size of this pool per replica (PGRST_DB_POOL) deliberately differs depending on the project level, verified in k8s_tenant.rs, the function that builds the PostgREST manifest for each project:
| Bearing | PGRST_DB_POOL / replica | Replicas | Connections / awakened project |
|---|---|---|---|
| Dedicated (premium, A1) | 10 (PostgREST default) | 2 | 20 |
| Shared (fleet, free/pro/team) | 2 (Aurabase default, lowered) | 2 | 4 |
On the dedicated level, the constraint is relaxed: a project has its own CNPG cluster, therefore its own max_connections, with no neighbors to spare. On the shared level, several projects from the same organization share a single cluster: it is this context which makes the connections budget decisive, developed in the following section.
The connections budget decides how many tenants are running at the same time
On a shared Postgres cluster, it is not the HTTP throughput that limits the number of simultaneously active projects. This is the number of connections their PostgREST instances hold open on the primary, compared to the available max_connections.
Aurabase derives this budget directly from the actual cluster limits, checked in fleet.rs: (max_connections − réserve superuser/CNPG − connexions réservées au pooler) ÷ connexions par projet éveillé, floor at 1. The fixed reserve is 10 connections (superuser, CNPG instance manager, metrics exporter, provisioner admin margin). On the defects delivered (pool of 2 per replica, 2 replicas, or 4 connections per awakened project), the calculation gives three different budgets depending on the sizing level of the cluster.
Source: derived from fleet.rs::derive_wake_budget and wake_budget_for_org_plan, Aurabase code, reread on August 24, 2026.
This budget is not a quota of owned projects: an team organization can hold 50 projects, most of which are dormant. This is a cap of concurrency: the number of projects that can hold open connections at the same time on the primary. A wake-up over budget does not fail, it is deferred until a sibling project goes back to sleep, checked in the same file. The topic of sizing max_connections itself is expanded upon in our article on tuning of max_connections, and the dedicated/mutualized tradeoff as a whole in dedicated vs. shared base.
Why Prefer: count=exact slows down a query on a large table
Asking for an exact total forces Postgres to count the visible rows of the filtered result with each query, a cost that grows with the table, not a free operation.
PostgreSQL does not maintain any indexed row counters out of the box. Under MVCC, the visibility of a row depends on the transaction that reads it. An exact COUNT(*) must therefore visit the candidate rows rather than reading a precomputed value. This is a well-documented structural limitation in the Postgres ecosystem, including analytics vendors like ClickHouse, which compare their own approximate counters to Postgres transactional behavior.
PostgREST documents these three counting strategies natively (postgrest.org, accessed August 24, 2026). The exact strategy guarantees a total at the price of the scan. planned returns an almost free estimate from the query planner. estimated automatically switches between the two based on a threshold. The choice is not cosmetic: a pagination which requires count=exact on a table of several million rows pays for this scan on each page, even when the user never consults the last one.
Truncation of db-max-rows is invisible without count=exact
A row cap can truncate a PostgREST response without any indication in the body or headers, unless explicitly requesting an exact total. We measured it in real conditions on a dedicated Aurabase instance, not assumed.
On a 10-row test table with PGRST_DB_MAX_ROWS=5, PostgREST v12.2.3 renders exactly the same Content-Range header for two very different situations:
| Query | Rendered lines | Content-Range | meta (Aurabase) |
|---|---|---|---|
| ?limit=50 (without count) | 5/10 real | 0-4/* | {} |
| ?limit=50&count=exact | 5/10 real | 0-4/10 | {total: 10} |
Without count=exact, the response of 5 rows is indistinguishable from a table which would really only contain 5: Content-Range: 0-4/* describes the rows rendered, never the limit applied. The actual cap in effect does not appear anywhere in this case, measured directly on the PostgREST path of the Aurabase SDK.
If your PostgREST deployment sets a db-max-rows (Aurabase defaults to 1000), a client that compares data.length to the requested limit to detect a full page may be mistaken. The error appears as soon as the server ceiling is lower than this limit. The only reliable signal is to compare the number of lines received to the total returned by count=exact, which directly brings into play the cost trade-off described in the previous section.
Reloading the schema cache after a migration
PostgREST keeps the Postgres schema in memory on startup. After a DDL (create table, add column), this cache must be reloaded before the new route responds, and this reload is asynchronous.
A write that arrives in this window may receive a transient 404 (cache not yet up to date), even though the table does indeed exist on the Postgres side. The Aurabase gateway absorbs this with a bounded retry loop, verified in postgrest_proxy.rs: up to 8 attempts, increasing backoff (250 ms plus 100 ms per attempt), 3.5 seconds cumulative in the worst case. This mechanism only affects writes, never reads.
The gateway does not emit any reload signal: it only waits. The only real trigger is an pg_notify('pgrst', 'reload schema') issued by the database service on the DDL path. If a migration path forgets to emit this signal, the 8 attempts exhaust on a cache that will never change, a risk documented as is in the code comment, not disguised.
For a self-hosted PostgREST deployment, the lesson generalizes. Each DDL path in your application should trigger the reload, via NOTIFY or an SIGUSR1 signal to the process. Otherwise, a migration produces a p99 latency spike disguised as intermittent errors right after deployment.
What the architecture slices, not the raw throughput
The four limitations documented here share one thing in common: none are seen on an isolated HTTP throughput test, yet all four determine whether a PostgREST deployment scales to production.
- Connection budget: caps the number of tenants active simultaneously on a shared cluster, regardless of the throughput per tenant.
- Cost of exact COUNT: grows with the table, not with the load; is bypassed with
planned/estimated. - Silent truncation: a correctly configured row cap can still break poorly instrumented paging.
- Schema reload: a latency window after each migration, bounded if the reload signal is well wired, unlimited otherwise.
Whether you're choosing between self-hosted PostgREST, a Hasura-style GraphQL layer, or a custom API, these four axes are a better point of comparison than an isolated req/s figure. See our comparison PostgREST vs Hasura vs custom API. The choice of the pooler that stands in front of your database matters just as much: our comparison PgBouncer vs Supavisor vs PgCat details why PostgREST cannot go through a pooler in transaction mode.
FAQs
What to remember
PostgREST almost never breaks under HTTP load alone: its architecture is too simple for that. What breaks in production is what surrounds it: how many connections its replicas keep open, how much an exact total costs. This also includes whether a truncation remains visible, and how long the window lasts after a migration.
These four limitations are not specific to Aurabase: they apply to any PostgREST deployment, self-hosted or managed. What this code shows is how a multi-tenant deployment makes them explicit rather than leaving them surprising in production.