PRODSovereign European BaaS platformOpen Dashboard →

Native AI · 10 min read

AI Gateway: native vs OpenAI-compatible providers

Affane Daylami · Fondateur · March 31, 2026

Back to blog

An AI Gateway routes calls to multiple LLM providers from a single entry point. At Aurabase, this layer lives directly in the Postgres backend. OpenAI, Anthropic (Claude) and Google Gemini each have a hard-coded native client, with its own error handling, token billing and streaming. Everything else, Mistral, Scaleway AI, self-hosted Ollama, goes through a generic OpenAI compatible adapter. The distinction is not cosmetic: it determines what really works (automatic switching, precise counting of reasoning tokens) and what works “as long as the provider faithfully imitates the OpenAI API”.

This English text was generated automatically from the French original and has not been reviewed yet.

We verified this distinction directly in the code of the aura-aiservice, not in a marketing page: three provider modules (anthropic, gemini, openai) make up the gateway, nothing more. The rest of the landscape confirms that this is a real developer category, not an isolated marketing argument. Neon publishes two dedicated pages (“AI Gateway” and “Backend for AI agents”), LiteLLM has established itself as a reference open source project, and Braintrust devotes its own comparisons to it.

The essentials

  • 3 native providers verified in code: OpenAI, Anthropic (Claude), Google Gemini (aura-ai/src/llm/mod.rs).
  • Mistral, Scaleway AI and Ollama go through the generic OpenAI adapter (OPENAI_BASE_URL), not through a dedicated client.
  • The native brings more than the connection: circuit breaker by (project, supplier), retry to backoff, fallback chain (default order Anthropic → OpenAI → Google), standardization of shutdown reasons.
  • Anthropic only covers chat: no embeddings API, unlike OpenAI and Gemini which cover both.
  • Neon, LiteLLM and Braintrust confirm the same category on the market side, with three different approaches: managed gateway, open source proxy, comparative content.
#
Definition

What is an AI Gateway on a Postgres backend?

An AI Gateway centralizes calls to external LLM providers behind a single interface, instead of coding each SDK integration on the application side. API keys remain server-side, never exposed to the client. The gateway adds a common layer of retry, failover and cost counting on top of providers with different response formats.

On a Postgres backend like Aurabase, this choice has a direct consequence: the same gateway powers the application chat, the NL2SQL (natural language translation to SQL) and the RAG (pgvector vector search). A poorly integrated provider degrades all three features at once, not just one. This is what makes the native/compatible distinction more than an implementation detail.

#
Technical distinction

Native provider or compatible endpoint: the concrete difference

A native client encodes the real form of the provider's API: request structure, response format, usage fields specific to this provider. This is the case of Anthropic, whose “Messages” API does not resemble that of OpenAI, or of Gemini, whose counting of reasoning tokens (thoughtsTokenCount) is added to the output counter instead of being already included there.

An OpenAI compatible endpoint reuses the existing OpenAI client and only changes the base URL. This works because the third-party provider (Mistral, Scaleway AI, Ollama) chose to imitate OpenAI's API contract, often with deviations: no separate reasoning token field, no guarantee on the exact form of errors. Compatibility ends where imitation ends.

#
Checked in code

3 native LLM clients at Aurabase, no more

The file that organizes the providers in aura-ai leaves no room for ambiguity. Three modules, one per native provider, nothing else declared.

llm/mod.rsrust
pub mod anthropic;
pub mod gemini;
pub mod openai;

Each module implements the ChatProvider trait (completion, streaming, model name). Two of them, OpenAI and Gemini, additionally implement EmbeddingProvider. Anthropic doesn't need it: Claude doesn't expose a vendor-side embeddings API, a product fact of Anthropic itself, not a shortcoming of the Aurabase code.

OpenAINATIVE CUSTOMERChat + calculation of high fidelity vector embeddings
Anthropic (Claude)NATIVE CUSTOMERChat inference & structured model completion Claude
Google GeminiNATIVE CUSTOMERChat + embeddings, additive reasoning token counting
MistralCOMPATIBLE LOCATIONRouting via standard OpenAI compatible protocol (custom URL)
Scaleway AICOMPATIBLE LOCATIONRoute via openai.rs client, variable OPENAI_BASE_URL
Ollama (self-hosted)COMPATIBLE LOCATIONRoute via openai.rs client, variable OPENAI_BASE_URL
#
Setup

Connect Mistral, Scaleway AI or Ollama to an Aurabase project

Configuring Mistral, Scaleway AI or Ollama does not require a new module: the same OPENAI_BASE_URL variable redirects the openai.rs client to another compatible endpoint. This is a configuration toggle, not development.

.env aura-aibash
# Default Provider: Native OpenAI
OPENAI_API_KEY=sk-...

# Switch to an OpenAI compatible provider (Mistral, Scaleway AI, Ollama...)
OPENAI_API_KEY=<clé du fournisseur choisi>
OPENAI_BASE_URL=<URL de base compatible OpenAI du fournisseur>
# e.g. Mistral: https://api.mistral.ai/v1

Behavior changes accordingly. HTTP errors remain classified by the same mechanism (429 → rate limit, 5xx → transient and retryable, 404 → unknown model), because the classification lives at the HTTP transport level, not provider-specific parsing. What doesn't follow: the precise counting of reasoning tokens, specific to the dedicated Gemini client.

#
Resilience

Why native is a game changer: switchover, errors, billing

The Aurabase gateway adds three resilience mechanisms on top of the three native clients. A circuit breaker per pair (project, provider) cuts calls to a repeatedly failed provider, with a probe token in a semi-open state before reopening it. An exponential backoff retry with jitter restarts transient errors (timeout, 5xx, 429), without dependence on an external random number library.

Retry and fallback naively do not stack

When several providers are configured in a chain, only one attempt is made per provider before switching to the next, to avoid amplification (retry × fallback) which would multiply upstream calls and total latency. The default order of this channel is Anthropic, then OpenAI, then Google Gemini.

Each provider also names the reason for stopping a response differently: length at OpenAI, MAX_TOKENS at Gemini, max_tokens at Anthropic, for the same reality (truncation). The code normalizes these three vocabularies towards a common set (stop, length, content_filter, tool_use, other). Without this standardization, a multi-vendor client would need to know all three vocabularies to detect a truncated response.

Billing illustrates the same risk. At OpenAI and Anthropic, the reasoning of the model is already included in the counter of output tokens charged. At Gemini, thoughtsTokenCount is added separately to candidatesTokenCount: ignoring it underestimates the real cost of a query. A generic OpenAI compatible adapter has no reason to know this particularity, specific to Gemini's native response format.

#
Landscape 2026

Neon, LiteLLM, Braintrust: where are the best LLM gateways in 2026?

The market confirms that an AI Gateway has become an expected brick, not an isolated marketing argument. Neon publishes two dedicated product pages, “AI Gateway” and “Backend for AI agents”, both Postgres developer oriented. LiteLLM has established itself as a reference open source project to unify calls to a large number of providers behind a format close to OpenAI. Braintrust, for its part, publishes its own comparisons on the subject, a sign that the category is strong enough to justify dedicated editorial content.

These players respond to a real need: reducing the coupling between the application code and a given LLM provider. The difference with Aurabase is the integration. The gateway does not live next to the backend: it shares the same service as NL2SQL and RAG, on the same Postgres database. The opposite compromise also exists: a dedicated proxy like LiteLLM generally covers more providers than a gateway integrated into an application backend.

AurabaseIntegrated with Postgres backend (aura-ai service)3 verified natives + OpenAI compatible for the rest
Neon AI GatewayDedicated product, alongside the managed Postgres databaseDocumented on two separate official pages
LiteLLMIndependent open source proxy, in front of any backendWide range of providers via a format close to OpenAI
#
Editorial honesty

When to choose a native gateway, when to choose a general proxy

A native gateway like that of Aurabase has a real advantage when the backend and the AI must remain in the same system: NL2SQL, RAG and application chat then share the same resilience policy and the same billing, without additional services to operate.

The opposite compromise exists. If your priority is to cover a very large number of providers, or if the gateway must serve several independent backends and not just a Postgres project, a general proxy like LiteLLM often remains the right choice. Aurabase does not seek to compete with this breadth of coverage: the bet is the depth across 3 major providers, integrated with the rest of the backend.

To understand how this integration concretely changes the use of NL2SQL compared to an approach using external connectors, the approach chosen by Supabase, see Supabase relies on connectors, not on native NL2SQL.

#
Overview

Integrated native gateway versus general LLM proxy

Summary of the criteria that really distinguish the two approaches, without value judgment: each responds to a different need.

SuppliersDepth on 3 major suppliers + OpenAI compatible for the restWide range of suppliers, generally uniform integration
API keysEncrypted on the backend side, same service as the databaseEncrypted on the proxy side, service separated from the application backend
NL2SQL / RAG linkSame service, same provider resolverNo native links, build-your-own integration
ResilienceCircuit breaker by (project, supplier), fallback, retry to backoffDepends on the configuration chosen for the proxy
DeploymentOne less service to operate (already in the backend)Detachable, reusable across multiple projects/backends

To see this gateway at work in a concrete case, see the NL2SQL tutorial on Postgres. For details of Aurabase's native AI capabilities, see the Native AIpage.

#
Frequently Asked Questions

FAQs

Can Mistral be used with Aurabase?+
Yes, via the OpenAI compatible endpoint: configure OPENAI_API_KEY with the Mistral key and OPENAI_BASE_URL with the Mistral API base URL. It is not a dedicated native client: Mistral uses the same code as the OpenAI provider, with the same limitations, including the absence of separate counting of reasoning tokens.
What happens if the primary LLM provider goes down?+
A breaker circuit per pair (project, supplier) detects repeated failures and cuts calls to this supplier. If a fallback chain is configured (default order: Anthropic, then OpenAI, then Google Gemini), the query automatically switches to the next provider, with only one attempt per provider to avoid retry × fallback amplification.
What is the difference between a native client and an OpenAI compatible endpoint?+
A native client encodes the real form of the provider's API: request structure, response format, usage fields specific to that provider, such as additive counting of reasoning tokens at Gemini. A compatible endpoint reuses the existing OpenAI client by changing only the base URL, which works as long as the third-party provider closely mimics the OpenAI API contract.
Does Aurabase offer embeddings for all native providers?+
No. OpenAI and Google Gemini implement the embeddings interface, not Anthropic. Claude does not expose a provider-side embeddings API: this is not a shortcoming of the Aurabase code, but a feature of the Anthropic product itself. The AI ​​Gateway technical guide details configuration by vendor, native or compatible.

READY TO DEPLOY?

Your backend in five minutes.

No credit card required · 500 MB free · 50,000 MAU