We verified this distinction directly in the code of the aura-aiservice, not in a marketing page: three provider modules (anthropic, gemini, openai) make up the gateway, nothing more. The rest of the landscape confirms that this is a real developer category, not an isolated marketing argument. Neon publishes two dedicated pages (“AI Gateway” and “Backend for AI agents”), LiteLLM has established itself as a reference open source project, and Braintrust devotes its own comparisons to it.
The essentials
- 3 native providers verified in code: OpenAI, Anthropic (Claude), Google Gemini (
aura-ai/src/llm/mod.rs). - Mistral, Scaleway AI and Ollama go through the generic OpenAI adapter (
OPENAI_BASE_URL), not through a dedicated client. - The native brings more than the connection: circuit breaker by (project, supplier), retry to backoff, fallback chain (default order Anthropic → OpenAI → Google), standardization of shutdown reasons.
- Anthropic only covers chat: no embeddings API, unlike OpenAI and Gemini which cover both.
- Neon, LiteLLM and Braintrust confirm the same category on the market side, with three different approaches: managed gateway, open source proxy, comparative content.
What is an AI Gateway on a Postgres backend?
An AI Gateway centralizes calls to external LLM providers behind a single interface, instead of coding each SDK integration on the application side. API keys remain server-side, never exposed to the client. The gateway adds a common layer of retry, failover and cost counting on top of providers with different response formats.
On a Postgres backend like Aurabase, this choice has a direct consequence: the same gateway powers the application chat, the NL2SQL (natural language translation to SQL) and the RAG (pgvector vector search). A poorly integrated provider degrades all three features at once, not just one. This is what makes the native/compatible distinction more than an implementation detail.
Native provider or compatible endpoint: the concrete difference
A native client encodes the real form of the provider's API: request structure, response format, usage fields specific to this provider. This is the case of Anthropic, whose “Messages” API does not resemble that of OpenAI, or of Gemini, whose counting of reasoning tokens (thoughtsTokenCount) is added to the output counter instead of being already included there.
An OpenAI compatible endpoint reuses the existing OpenAI client and only changes the base URL. This works because the third-party provider (Mistral, Scaleway AI, Ollama) chose to imitate OpenAI's API contract, often with deviations: no separate reasoning token field, no guarantee on the exact form of errors. Compatibility ends where imitation ends.
3 native LLM clients at Aurabase, no more
The file that organizes the providers in aura-ai leaves no room for ambiguity. Three modules, one per native provider, nothing else declared.
Each module implements the ChatProvider trait (completion, streaming, model name). Two of them, OpenAI and Gemini, additionally implement EmbeddingProvider. Anthropic doesn't need it: Claude doesn't expose a vendor-side embeddings API, a product fact of Anthropic itself, not a shortcoming of the Aurabase code.
| OpenAI | NATIVE CUSTOMER | Chat + calculation of high fidelity vector embeddings |
|---|---|---|
| Anthropic (Claude) | NATIVE CUSTOMER | Chat inference & structured model completion Claude |
| Google Gemini | NATIVE CUSTOMER | Chat + embeddings, additive reasoning token counting |
| Mistral | COMPATIBLE LOCATION | Routing via standard OpenAI compatible protocol (custom URL) |
| Scaleway AI | COMPATIBLE LOCATION | Route via openai.rs client, variable OPENAI_BASE_URL |
| Ollama (self-hosted) | COMPATIBLE LOCATION | Route via openai.rs client, variable OPENAI_BASE_URL |
Connect Mistral, Scaleway AI or Ollama to an Aurabase project
Configuring Mistral, Scaleway AI or Ollama does not require a new module: the same OPENAI_BASE_URL variable redirects the openai.rs client to another compatible endpoint. This is a configuration toggle, not development.
Behavior changes accordingly. HTTP errors remain classified by the same mechanism (429 → rate limit, 5xx → transient and retryable, 404 → unknown model), because the classification lives at the HTTP transport level, not provider-specific parsing. What doesn't follow: the precise counting of reasoning tokens, specific to the dedicated Gemini client.
Why native is a game changer: switchover, errors, billing
The Aurabase gateway adds three resilience mechanisms on top of the three native clients. A circuit breaker per pair (project, provider) cuts calls to a repeatedly failed provider, with a probe token in a semi-open state before reopening it. An exponential backoff retry with jitter restarts transient errors (timeout, 5xx, 429), without dependence on an external random number library.
When several providers are configured in a chain, only one attempt is made per provider before switching to the next, to avoid amplification (retry × fallback) which would multiply upstream calls and total latency. The default order of this channel is Anthropic, then OpenAI, then Google Gemini.
Each provider also names the reason for stopping a response differently: length at OpenAI, MAX_TOKENS at Gemini, max_tokens at Anthropic, for the same reality (truncation). The code normalizes these three vocabularies towards a common set (stop, length, content_filter, tool_use, other). Without this standardization, a multi-vendor client would need to know all three vocabularies to detect a truncated response.
Billing illustrates the same risk. At OpenAI and Anthropic, the reasoning of the model is already included in the counter of output tokens charged. At Gemini, thoughtsTokenCount is added separately to candidatesTokenCount: ignoring it underestimates the real cost of a query. A generic OpenAI compatible adapter has no reason to know this particularity, specific to Gemini's native response format.
Neon, LiteLLM, Braintrust: where are the best LLM gateways in 2026?
The market confirms that an AI Gateway has become an expected brick, not an isolated marketing argument. Neon publishes two dedicated product pages, “AI Gateway” and “Backend for AI agents”, both Postgres developer oriented. LiteLLM has established itself as a reference open source project to unify calls to a large number of providers behind a format close to OpenAI. Braintrust, for its part, publishes its own comparisons on the subject, a sign that the category is strong enough to justify dedicated editorial content.
These players respond to a real need: reducing the coupling between the application code and a given LLM provider. The difference with Aurabase is the integration. The gateway does not live next to the backend: it shares the same service as NL2SQL and RAG, on the same Postgres database. The opposite compromise also exists: a dedicated proxy like LiteLLM generally covers more providers than a gateway integrated into an application backend.
| Aurabase | Integrated with Postgres backend (aura-ai service) | 3 verified natives + OpenAI compatible for the rest |
|---|---|---|
| Neon AI Gateway | Dedicated product, alongside the managed Postgres database | Documented on two separate official pages |
| LiteLLM | Independent open source proxy, in front of any backend | Wide range of providers via a format close to OpenAI |
When to choose a native gateway, when to choose a general proxy
A native gateway like that of Aurabase has a real advantage when the backend and the AI must remain in the same system: NL2SQL, RAG and application chat then share the same resilience policy and the same billing, without additional services to operate.
The opposite compromise exists. If your priority is to cover a very large number of providers, or if the gateway must serve several independent backends and not just a Postgres project, a general proxy like LiteLLM often remains the right choice. Aurabase does not seek to compete with this breadth of coverage: the bet is the depth across 3 major providers, integrated with the rest of the backend.
To understand how this integration concretely changes the use of NL2SQL compared to an approach using external connectors, the approach chosen by Supabase, see Supabase relies on connectors, not on native NL2SQL.
Integrated native gateway versus general LLM proxy
Summary of the criteria that really distinguish the two approaches, without value judgment: each responds to a different need.
| Suppliers | Depth on 3 major suppliers + OpenAI compatible for the rest | Wide range of suppliers, generally uniform integration |
|---|---|---|
| API keys | Encrypted on the backend side, same service as the database | Encrypted on the proxy side, service separated from the application backend |
| NL2SQL / RAG link | Same service, same provider resolver | No native links, build-your-own integration |
| Resilience | Circuit breaker by (project, supplier), fallback, retry to backoff | Depends on the configuration chosen for the proxy |
| Deployment | One less service to operate (already in the backend) | Detachable, reusable across multiple projects/backends |
To see this gateway at work in a concrete case, see the NL2SQL tutorial on Postgres. For details of Aurabase's native AI capabilities, see the Native AIpage.