PRODSovereign European BaaS platformOpen Dashboard →

Native AI · 7 min read

Embedding size: 768, 1536 or 3072 dimensions?

Affane Daylami · Fondateur · April 3, 2026

Back to blog

The choice between 768, 1536 and 3072 dimensions is not a cosmetic adjustment. It sets the storage volume of your vector database, the type of index that can be used and the price paid for each API call. OpenAI itself documents the quality gap between its two current models: text-embedding-3-small, in 1536 native dimensions, reached 62.3% on the MTEB benchmark, compared to 64.6% for text-embedding-3-large in 3072 native dimensions. A real gain, but one that is paid for elsewhere than you might think.

This English text was generated automatically from the French original and has not been reviewed yet.

This article is based on the official documentation published by OpenAI, on team feedback collected in OpenAI developer community discussions, and on the behavior verified in the Aurabase code, which natively routes these three dimension classes to three distinct vector columns. Each external figure is dated and sourced; for the methodology we apply to our own measurements, see our benchmark methodology pillar.

The essentials
  • 1536 dimensions remains the most balanced choice for the majority of cases: native text-embedding-3-small, or truncated text-embedding-3-large, without departing from pgvector's native HNSW support.
  • 3072 dimensions (text-embedding-3-large) offers the highest MTEB score published by OpenAI (64.6% vs. 62.3%), but exceeds the 2000 dimension limit of pgvector's vector type: the HNSW index requires a cast to halfvec.
  • Truncating an embedding via the OpenAI dimensions parameter (Matryoshka technique) reduces storage and speeds up the search, but does not reduce the price: this depends on the model queried, not on the size of the vector returned.
  • Thanks to halfvec storage (2 bytes per dimension), a 3072-dimensional vector occupies the same raw disk space at Aurabase as a 1536 vector in classic vector (4 bytes per dimension): approximately 6 KB.
  • Aurabase natively supports exactly 3 dimension classes, 768, 1536 and 3072, each on its own column (checked in aura-ai/src/embeddings/mod.rs): no free dimension field.
#
Overview

768, 1536 or 3072: what each level really changes

Clearly, choosing between text-embedding-3-small and text-embedding-3-large first amounts to choosing between 1536 and 3072 native dimensions, before even talking about truncation. The table below summarizes the verifiable facts on the three levels that pgvector natively recognizes on the indexing side, and that Aurabase route.

Criterion768 dimensions1536 dimensions3072 dimensions
Related model(s)OpenAI truncation, or legacy/open source (native) modeltext-embedding-3-small (native) or truncated 3-largetext-embedding-3-large (native)
Average MTEB scorenot natively released by OpenAI at this size62,3 %64,6 %
Indicative OpenAI price / 1M tokensdepends on the model queried, not on the size$0.02 (small) or $0.13 (large truncated)$0.13 (text-embedding-3-large)
Gross stored weight / vector3 KB (float32)6 KB (float32)12 KB (vector) or 6 KB (halfvec, Aurabase)
Native HNSW pgvector indexyesyesno: cast halfvec required (>2000 dims)
Aurabase column (verified code)embedding_768embedding_1536embedding_3072

Sources: OpenAI, official blog “New embedding models and API updates”, January 25, 2024 (MTEB scores and launch prices, check on the current pricing page before use); aura-ai/src/embeddings/mod.rs, Aurabase (columns and index support, verified August 24, 2026).

#
Storage

The impact on storage: the calculation that changes everything

An embedding is stored like an array of floating point numbers. In pgvector, the classic vector type encodes each dimension on 4 bytes (float32): 768 dimensions therefore weigh around 3 KB of raw data per vector, 1536 dimensions around 6 KB, and 3072 dimensions around 12 KB, even before counting the pgvector header and the Postgres page overhead.

This is where the halfvec type of pgvector comes in, which encodes each dimension in 2 bytes (float16) instead of 4. A 3072-dimensional vector stored in halfvec weighs around 6 KB: exactly the weight of a 1536-dimensional vector stored in classic vector.

3 KB
768 sun
vector, float32 (4 bytes/dim)
6 KB
1536 Sun
vector, float32: same weight as 3072 in halfvec
12 KB
3072 sun
classic vector, float32 (before cast halfvec)

Direct and unintuitive consequence: at Aurabase, going from 1536 to 3072 dimensions does not double the actual storage on disk, since the embedding_3072 column is queried via a halfveccast. The real additional cost of 3072 dimensions is therefore not primarily the disk: it is the price of the associated OpenAI model, and the output of native index support of the vectortype, detailed in the following section.

#
pgvector / HNSW

Why 3072 dimensions change index type under pgvector

The vector type of pgvector does not allow building an HNSW or IVFFlat index beyond 2000 dimensions. 3072 dimensions therefore exceeds this limit: no vector search query on a vector(3072) column can rely on an approximate index, it falls back on a complete sequential scan, unusable on the scale of a RAG corpus in production.

The Aurabase code handles this case explicitly: the embedding_3072 column is cast to halfvec(3072) on each insert and search query, a type that pgvector can index up to 4000 dimensions. Columns 768 and 1536 remain native vector, uncast, since they do not approach the limit.

embeddings/mod.rs (extrait simplifié)rust
// For 3072 (>2000), HNSW does not index type `vector` → cast halfvec
fn vec_query_parts(dims: usize) -> AuraResult<(&'static str, &'static str, &'static str)> {
    match dims {
        768  => Ok(("embedding_768",  "", "::vector")),
        1536 => Ok(("embedding_1536", "", "::vector")),
        3072 => Ok(("embedding_3072", "::halfvec(3072)", "::halfvec(3072)")),
        other => Err(...), // unsupported dimension
    }
}

This detail also explains why an embedding dimension not listed in [768, 1536, 3072] explicitly fails on the Aurabase side, rather than being accepted and then poorly indexed: the column name always comes from a fixed allowlist, never from a free value sent by the client. To dig deeper into building an HNSW index on Postgres beyond this specific case, see our article HNSW index and Postgres vector search.

#
API cost

Reducing dimension without losing everything: OpenAI’s Matryoshka truncation

As of January 2024, OpenAI's Embeddings API accepts a dimensions parameter that shortens the returned vector without re-invoking a different model. The technique is called Matryoshka Representation Learning: the model is trained to concentrate useful information in the first dimensions of the vector, so that a truncation loses precision gradually rather than abruptly.

OpenAI illustrates the effectiveness of this technique with a specific example in its announcement: text-embedding-3-large, truncated to only 256 dimensions, still exceeds the MTEB score of the old text-embedding-ada-002 used at its full size of 1536 dimensions (source: OpenAI, official blog, January 25, 2024). A vector 12 times smaller which performs better than a full vector, on this specific benchmark.

llm/openai.rs (extrait, requête d'embedding)rust
let body = json!({
    "model": self.embed_model,       // e.g. "text-embedding-3-large"
    "input": text,
    "dimensions": self.embed_dimensions, // truncates 3072 → the configured value
});

Important point, and often misunderstood: truncating does not reduce the price charged. OpenAI charges based on the model queried, not the size of the vector returned, since the actual cost is the calculation performed on the input text. Requesting 1536 dimensions from text-embedding-3-large therefore costs the same price as its 3072 native dimensions (source: OpenAI, official blog, January 25, 2024); only the storage and search speed change.

This is precisely the default choice of Aurabase, verified in config/mod.rs: the model configured by default is text-embedding-3-large, but the output dimension configured by default is 1536, not 3072. The service therefore pays for the representation of the wide model, truncated to remain on an indexable vector column in native HNSW, without the halfvec cast required to 3072.

#
Decision

Which dimension to choose according to your use case

1536 dimensions remains the reasonable starting point for the majority of RAG or semantic search projects: the MTEB score of text-embedding-3-small (62.3%) remains close to that of the large model, storage remains light, and the classic vector type of pgvector indexes in HNSW without any particular configuration.

3072 dimensions is justified when the corpus is ambiguous or technical, where the quality gap between 62.3% and 64.6% translates into visibly better search results on your own queries, not on the general OpenAI benchmark. Several team feedbacks recorded in the OpenAI developer community discussions point in this direction: the gain of 3072 dimensions is measured on a case-by-case basis, it cannot be assumed.

768 dimensions is especially suitable when volume takes precedence over nuance: a large corpus where the storage or calculation budget is the real constraint, or the use of a legacy embedding model already in 768 native dimensions.

A simple rule before deciding

Do not set the dimension until you have measured search quality on a representative sample of your own corpus, not just the general MTEB score published by OpenAI. The MTEB averages dozens of heterogeneous tasks; your RAG corpus is just one.

This choice of dimension is part of a larger RAG stack, embeddings, HNSW index, hybrid search, which our page Native AIdocuments.

#
Frequently asked questions

What we get asked most often

Can we change the size of an already indexed corpus without reindexing everything?+
No. Aurabase filters each semantic search by exact model AND dimension (embedding_model column + column dedicated to the dimension). A corpus indexed in 1536 dimensions becomes invisible to a search carried out in 3072, and vice versa: changing dimension requires reindexing the corpus under the new class.
Does Gemini allow you to go beyond 3072 dimensions?+
No. Aurabase's Gemini client explicitly caps the dimension at 3072 (outputDimensionality). 3072 is the ceiling common to the three classes supported by Aurabase, all suppliers combined.
Should you always choose 3072 dimensions for best results?+
Not necessarily. The difference in MTEB score between 1536 and 3072 (62.3% versus 64.6% according to OpenAI) remains modest in the face of the architectural change implied by 3072: release of native HNSW support of the vectortype, mandatory halfvec cast, and price of the wide model. The gain must be verified on your corpus before justifying this cost.
text-embedding-3-small or text-embedding-3-large for a RAG project in production?+
It depends on the budget and the nature of the corpus, not a universal rule. Text-embedding-3-small (1536 native dimensions) covers the majority of cases at lower cost; text-embedding-3-large is justified on an ambiguous corpus where the gain in precision is measured concretely on your own test queries, not just on the general MTEB score.
#
In summary

The right choice is not the greatest, it is the best measured

768, 1536 and 3072 dimensions are not divided on a single axis. 3072 gains in MTEB score published by OpenAI, but leaves pgvector's native HNSW support and pays the price of the large model, whatever the dimension ultimately requested. 1536 remains the most common default balance, including at Aurabase. 768 serves cases where volume takes precedence over nuance.

OpenAI's dimensions parameter changes the question to ask: it is no longer "which model to choose", but "which truncation to accept, for what gain measured on my corpus". Before finalizing a choice in production, test the search quality on a real sample, not just on a general benchmark. Our guide RAG pipeline with pgvector details the complete setup, from ingestion to hybrid search.

READY TO DEPLOY?

Your backend in five minutes.

No credit card required · 500 MB free · 50,000 MAU