PRODSovereign European BaaS platformOpen Dashboard →

Performance · 11 min read

PostgREST: benchmark and real limits in production

Affane Daylami · Fondateur · May 18, 2026

Back to blog

PostgREST itself is almost never the bottleneck. On dedicated Aurabase instances, a replica runs with 50 to 250 millicores of CPU and 64 to 128 MB of RAM. It's a lightweight Haskell binary that translates HTTP requests to SQL, nothing more. The real limits that appear in production are elsewhere. Four of them come up most often: the budget for Postgres connections that its replicas consume, and the cost of an exact COUNT under MVCC. A response truncation can also remain invisible in the headers, as can a latency window after each schema migration.

This English text was generated automatically from the French original and has not been reviewed yet.

Our article on PostgREST compatibility details what the server covers functionally (filters, embedding, RPC, RLS) and what it leaves up to you. This one comes from somewhere else. It documents, with Aurabase code and the official PostgREST documentation as sources, where and why PostgREST actually plateaus at scale. We are not reproducing here a load bank that we have not run ourselves. Our benchmark methodology explains why an isolated figure, without a published protocol, does not seem reliable to us.

The essentials

  • PostgREST itself is lightweight: 50 to 250 millicores of CPU, 64 to 128 MB of RAM per replica on dedicated Aurabase instances. Raw HTTP throughput is almost never the limiting factor in production.
  • The real ceiling is the Postgres connection budget: PGRST_DB_POOL × replicas. Verified in the Aurabase code: 20 connections per project on the dedicated level (10×2), 4 on the shared level (2×2). This is a deliberate choice to fit more tenants on the same max_connections.
  • Prefer: count=exact forces expensive MVCC scanning on large tables. PostgREST documents two cheaper alternatives: count=planned and count=estimated, costing an approximate total.
  • A db-max-rows ceiling (1000 lines by default at Aurabase) truncates a response WITHOUT reporting it in Content-Range (measured in real conditions, detailed below).
  • After a DDL migration, the PostgREST schema cache reloads asynchronously. The Aurabase gateway retries up to 8 times (around 3.5 seconds cumulative in the worst case) before giving up, a behavior documented directly in the code.
#
Methodology

What a PostgREST benchmark measures, and what it doesn't measure

An HTTP throughput test on PostgREST mainly measures Postgres, rarely PostgREST. The server is a thin layer of translation in front of the base. In the vast majority of real-world loads, response time is dominated by the SQL query executed, not the process that generated it.

The PostgREST project maintains a dedicated repository for this topic, PostgREST/postgrest-benchmark on GitHub, which tracks throughput variations from release to release rather than publishing an isolated marketing figure. We have neither performed nor republished it here. Its results depend on hardware, schematic size, and scenario tested, exactly the variables that our own benchmark protocol requires to be documented before quoting a figure.

Below PostgREST, it's pgbench which measures the layer that really matters: SQL transaction time under concurrent load. This is the official PostgreSQL benchmark tool (postgresql.org/docs/current/pgbench.html, accessed August 24, 2026). Rather than reproducing this protocol here, this article documents four concrete architectural limitations of PostgREST in production, each verified in the Aurabase source code or in the official project documentation.

#
Checked in code

The actual footprint of a PostgREST instance at Aurabase

Each Aurabase Postgres engine project receives two dedicated PostgREST replicas, co-located with its cluster. The Kubernetes manifest that deploys them sets modest resources.

50-250m
CPU PER REPLICA
requests → limits
64-128
MB RAM PER REPLICA
requests → limits
2
REPLICAS BY PROJECT
high availability (P22)

What these replicas really consume is not CPU: they are connections to the Postgres primary. Each PostgREST instance connects directly to the primary (-rw), without going through the PgBouncer pooler deployed for the tenant. This choice is already detailed in our article on PostgREST compatibility: the LISTEN/NOTIFY schema reloading mechanism requires a persistent connection, incompatible with a pooler in transaction mode. What this article adds: how much does it actually cost, in terms of connections, and where it peaks.

The size of this pool per replica (PGRST_DB_POOL) deliberately differs depending on the project level, verified in k8s_tenant.rs, the function that builds the PostgREST manifest for each project:

BearingPGRST_DB_POOL / replicaReplicasConnections / awakened project
Dedicated (premium, A1)10 (PostgREST default)220
Shared (fleet, free/pro/team)2 (Aurabase default, lowered)24
deploy/cnpg/tenant-postgrest.yaml (real extract, value substituted by the provisioner)yaml
# Fingerprint of connections per replica on the primary.
env:
  - { name: PGRST_DB_POOL, value: "PGRST_DB_POOL_VALUE" }  # 10 (dedicated) or 2 (shared)

On the dedicated level, the constraint is relaxed: a project has its own CNPG cluster, therefore its own max_connections, with no neighbors to spare. On the shared level, several projects from the same organization share a single cluster: it is this context which makes the connections budget decisive, developed in the following section.

#
The real ceiling

The connections budget decides how many tenants are running at the same time

On a shared Postgres cluster, it is not the HTTP throughput that limits the number of simultaneously active projects. This is the number of connections their PostgREST instances hold open on the primary, compared to the available max_connections.

Aurabase derives this budget directly from the actual cluster limits, checked in fleet.rs: (max_connections − réserve superuser/CNPG − connexions réservées au pooler) ÷ connexions par projet éveillé, floor at 1. The fixed reserve is 10 connections (superuser, CNPG instance manager, metrics exporter, provisioner admin margin). On the defects delivered (pool of 2 per replica, 2 replicas, or 4 connections per awakened project), the calculation gives three different budgets depending on the sizing level of the cluster.

Budget for simultaneously active projects by level, shared Postgres clusterFree level: 5 simultaneous active projects (max_connections 50, pooler 20). Pro level: 7 (max_connections 100, pooler 60). Team level: 10 (max_connections 200, pooler 150). Formula derived from Aurabase code (fleet.rs::derive_wake_budget), fixed reserve of 10 connections, 4 connections per awake project.024681012free (max_connections 50)5 projectspro (max_connections 100)7 projectsteam (max_connections 200)10 projects

Source: derived from fleet.rs::derive_wake_budget and wake_budget_for_org_plan, Aurabase code, reread on August 24, 2026.

This budget is not a quota of owned projects: an team organization can hold 50 projects, most of which are dormant. This is a cap of concurrency: the number of projects that can hold open connections at the same time on the primary. A wake-up over budget does not fail, it is deferred until a sibling project goes back to sleep, checked in the same file. The topic of sizing max_connections itself is expanded upon in our article on tuning of max_connections, and the dedicated/mutualized tradeoff as a whole in dedicated vs. shared base.

#
Hidden cost

Why Prefer: count=exact slows down a query on a large table

Asking for an exact total forces Postgres to count the visible rows of the filtered result with each query, a cost that grows with the table, not a free operation.

PostgreSQL does not maintain any indexed row counters out of the box. Under MVCC, the visibility of a row depends on the transaction that reads it. An exact COUNT(*) must therefore visit the candidate rows rather than reading a precomputed value. This is a well-documented structural limitation in the Postgres ecosystem, including analytics vendors like ClickHouse, which compare their own approximate counters to Postgres transactional behavior.

terminalbash
# Expensive on a large table: forces an MVCC scan of the filtered result
curl "https://<gateway>/v1/db/<project_id>/orders?status=eq.paid" \
  -H "apikey: <clé>" -H "Prefer: count=exact"

# Less expensive alternatives, documented by PostgREST
  -H "Prefer: count=planned"   # estimation via planner
  -H "Prefer: count=estimated" # planned beyond a threshold, exact below

PostgREST documents these three counting strategies natively (postgrest.org, accessed August 24, 2026). The exact strategy guarantees a total at the price of the scan. planned returns an almost free estimate from the query planner. estimated automatically switches between the two based on a threshold. The choice is not cosmetic: a pagination which requires count=exact on a table of several million rows pays for this scan on each page, even when the user never consults the last one.

#
Measured in real

Truncation of db-max-rows is invisible without count=exact

A row cap can truncate a PostgREST response without any indication in the body or headers, unless explicitly requesting an exact total. We measured it in real conditions on a dedicated Aurabase instance, not assumed.

On a 10-row test table with PGRST_DB_MAX_ROWS=5, PostgREST v12.2.3 renders exactly the same Content-Range header for two very different situations:

QueryRendered linesContent-Rangemeta (Aurabase)
?limit=50 (without count)5/10 real0-4/*{}
?limit=50&count=exact5/10 real0-4/10{total: 10}

Without count=exact, the response of 5 rows is indistinguishable from a table which would really only contain 5: Content-Range: 0-4/* describes the rows rendered, never the limit applied. The actual cap in effect does not appear anywhere in this case, measured directly on the PostgREST path of the Aurabase SDK.

Consequence for any pagination on PostgREST

If your PostgREST deployment sets a db-max-rows (Aurabase defaults to 1000), a client that compares data.length to the requested limit to detect a full page may be mistaken. The error appears as soon as the server ceiling is lower than this limit. The only reliable signal is to compare the number of lines received to the total returned by count=exact, which directly brings into play the cost trade-off described in the previous section.

#
Deferred latency

Reloading the schema cache after a migration

PostgREST keeps the Postgres schema in memory on startup. After a DDL (create table, add column), this cache must be reloaded before the new route responds, and this reload is asynchronous.

A write that arrives in this window may receive a transient 404 (cache not yet up to date), even though the table does indeed exist on the Postgres side. The Aurabase gateway absorbs this with a bounded retry loop, verified in postgrest_proxy.rs: up to 8 attempts, increasing backoff (250 ms plus 100 ms per attempt), 3.5 seconds cumulative in the worst case. This mechanism only affects writes, never reads.

A detail that the code itself documents

The gateway does not emit any reload signal: it only waits. The only real trigger is an pg_notify('pgrst', 'reload schema') issued by the database service on the DDL path. If a migration path forgets to emit this signal, the 8 attempts exhaust on a cache that will never change, a risk documented as is in the code comment, not disguised.

For a self-hosted PostgREST deployment, the lesson generalizes. Each DDL path in your application should trigger the reload, via NOTIFY or an SIGUSR1 signal to the process. Otherwise, a migration produces a p99 latency spike disguised as intermittent errors right after deployment.

#
Summary

What the architecture slices, not the raw throughput

The four limitations documented here share one thing in common: none are seen on an isolated HTTP throughput test, yet all four determine whether a PostgREST deployment scales to production.

  • Connection budget: caps the number of tenants active simultaneously on a shared cluster, regardless of the throughput per tenant.
  • Cost of exact COUNT: grows with the table, not with the load; is bypassed with planned/estimated.
  • Silent truncation: a correctly configured row cap can still break poorly instrumented paging.
  • Schema reload: a latency window after each migration, bounded if the reload signal is well wired, unlimited otherwise.

Whether you're choosing between self-hosted PostgREST, a Hasura-style GraphQL layer, or a custom API, these four axes are a better point of comparison than an isolated req/s figure. See our comparison PostgREST vs Hasura vs custom API. The choice of the pooler that stands in front of your database matters just as much: our comparison PgBouncer vs Supavisor vs PgCat details why PostgREST cannot go through a pooler in transaction mode.

#
Frequently Asked Questions

FAQs

Is PostgREST fast enough for large-scale production?+
PostgREST itself is a lightweight process. On dedicated Aurabase instances, a replica runs with 50 to 250 millicores of CPU and 64 to 128 MB of RAM, verified in the project's Kubernetes manifest. Raw HTTP throughput is almost never the limiting factor in production. It's the Postgres connection budget, the cost of an exact COUNT, and the schema cache that determine whether the whole thing scales, not the speed of the PostgREST binary alone.
How do I know if my PostgREST response was truncated by db-max-rows?+
The Content-Range header returned by PostgREST never says this. A response capped at 5 rows per db-max-rows is indistinguishable from a table which really only contains 5, measured in real conditions on a dedicated Aurabase instance. The only reliable way to detect it is to compare the number of lines received to the total returned by Prefer: count=exact, without this header the truncation remains invisible.
Does the exact COUNT always slow down a PostgREST request?+
Prefer: count=exact forces Postgres to count the visible rows of the filtered result with each query, a cost that increases with table size due to MVCC. Postgres does not maintain an indexed row counter out of the box. PostgREST offers two less expensive alternatives, count=planned (estimation via the scheduler) and count=estimated (automatic switching beyond a threshold), documented in its official documentation.
How many Postgres connections does PostgREST consume?+
It depends entirely on PGRST_DB_POOL multiplied by the number of replicas. Verified in the Aurabase code: a dedicated instance (premium tier) opens by default 10 connections per replica, or 20 in total over 2 replicas. The shared level voluntarily lowers this pool to 2 per replica, or 4 connections per awakened project, to accommodate more tenants on the same max_connections budget of the shared cluster.
Is there an official PostgREST benchmark?+
The project maintains a dedicated repository, PostgREST/postgrest-benchmark on GitHub, which tracks throughput variations from release to release rather than publishing an isolated marketing figure. We have neither performed nor republished it here. This article documents verified architectural limitations in our code and in the official PostgREST documentation, not a bench we reproduced ourselves.
#
Conclusion

What to remember

PostgREST almost never breaks under HTTP load alone: its architecture is too simple for that. What breaks in production is what surrounds it: how many connections its replicas keep open, how much an exact total costs. This also includes whether a truncation remains visible, and how long the window lasts after a migration.

These four limitations are not specific to Aurabase: they apply to any PostgREST deployment, self-hosted or managed. What this code shows is how a multi-tenant deployment makes them explicit rather than leaving them surprising in production.

READY TO DEPLOY?

Your backend in five minutes.

No credit card required · 500 MB free · 50,000 MAU