This article is based on the official documentation published by OpenAI, on team feedback collected in OpenAI developer community discussions, and on the behavior verified in the Aurabase code, which natively routes these three dimension classes to three distinct vector columns. Each external figure is dated and sourced; for the methodology we apply to our own measurements, see our benchmark methodology pillar.
- 1536 dimensions remains the most balanced choice for the majority of cases: native text-embedding-3-small, or truncated text-embedding-3-large, without departing from pgvector's native HNSW support.
- 3072 dimensions (text-embedding-3-large) offers the highest MTEB score published by OpenAI (64.6% vs. 62.3%), but exceeds the 2000 dimension limit of pgvector's
vectortype: the HNSW index requires a cast tohalfvec. - Truncating an embedding via the OpenAI
dimensionsparameter (Matryoshka technique) reduces storage and speeds up the search, but does not reduce the price: this depends on the model queried, not on the size of the vector returned. - Thanks to
halfvecstorage (2 bytes per dimension), a 3072-dimensional vector occupies the same raw disk space at Aurabase as a 1536 vector in classicvector(4 bytes per dimension): approximately 6 KB. - Aurabase natively supports exactly 3 dimension classes, 768, 1536 and 3072, each on its own column (checked in
aura-ai/src/embeddings/mod.rs): no free dimension field.
768, 1536 or 3072: what each level really changes
Clearly, choosing between text-embedding-3-small and text-embedding-3-large first amounts to choosing between 1536 and 3072 native dimensions, before even talking about truncation. The table below summarizes the verifiable facts on the three levels that pgvector natively recognizes on the indexing side, and that Aurabase route.
| Criterion | 768 dimensions | 1536 dimensions | 3072 dimensions |
|---|---|---|---|
| Related model(s) | OpenAI truncation, or legacy/open source (native) model | text-embedding-3-small (native) or truncated 3-large | text-embedding-3-large (native) |
| Average MTEB score | not natively released by OpenAI at this size | 62,3 % | 64,6 % |
| Indicative OpenAI price / 1M tokens | depends on the model queried, not on the size | $0.02 (small) or $0.13 (large truncated) | $0.13 (text-embedding-3-large) |
| Gross stored weight / vector | 3 KB (float32) | 6 KB (float32) | 12 KB (vector) or 6 KB (halfvec, Aurabase) |
| Native HNSW pgvector index | yes | yes | no: cast halfvec required (>2000 dims) |
| Aurabase column (verified code) | embedding_768 | embedding_1536 | embedding_3072 |
Sources: OpenAI, official blog “New embedding models and API updates”, January 25, 2024 (MTEB scores and launch prices, check on the current pricing page before use); aura-ai/src/embeddings/mod.rs, Aurabase (columns and index support, verified August 24, 2026).
The impact on storage: the calculation that changes everything
An embedding is stored like an array of floating point numbers. In pgvector, the classic vector type encodes each dimension on 4 bytes (float32): 768 dimensions therefore weigh around 3 KB of raw data per vector, 1536 dimensions around 6 KB, and 3072 dimensions around 12 KB, even before counting the pgvector header and the Postgres page overhead.
This is where the halfvec type of pgvector comes in, which encodes each dimension in 2 bytes (float16) instead of 4. A 3072-dimensional vector stored in halfvec weighs around 6 KB: exactly the weight of a 1536-dimensional vector stored in classic vector.
Direct and unintuitive consequence: at Aurabase, going from 1536 to 3072 dimensions does not double the actual storage on disk, since the embedding_3072 column is queried via a halfveccast. The real additional cost of 3072 dimensions is therefore not primarily the disk: it is the price of the associated OpenAI model, and the output of native index support of the vectortype, detailed in the following section.
Why 3072 dimensions change index type under pgvector
The vector type of pgvector does not allow building an HNSW or IVFFlat index beyond 2000 dimensions. 3072 dimensions therefore exceeds this limit: no vector search query on a vector(3072) column can rely on an approximate index, it falls back on a complete sequential scan, unusable on the scale of a RAG corpus in production.
The Aurabase code handles this case explicitly: the embedding_3072 column is cast to halfvec(3072) on each insert and search query, a type that pgvector can index up to 4000 dimensions. Columns 768 and 1536 remain native vector, uncast, since they do not approach the limit.
This detail also explains why an embedding dimension not listed in [768, 1536, 3072] explicitly fails on the Aurabase side, rather than being accepted and then poorly indexed: the column name always comes from a fixed allowlist, never from a free value sent by the client. To dig deeper into building an HNSW index on Postgres beyond this specific case, see our article HNSW index and Postgres vector search.
Reducing dimension without losing everything: OpenAI’s Matryoshka truncation
As of January 2024, OpenAI's Embeddings API accepts a dimensions parameter that shortens the returned vector without re-invoking a different model. The technique is called Matryoshka Representation Learning: the model is trained to concentrate useful information in the first dimensions of the vector, so that a truncation loses precision gradually rather than abruptly.
OpenAI illustrates the effectiveness of this technique with a specific example in its announcement: text-embedding-3-large, truncated to only 256 dimensions, still exceeds the MTEB score of the old text-embedding-ada-002 used at its full size of 1536 dimensions (source: OpenAI, official blog, January 25, 2024). A vector 12 times smaller which performs better than a full vector, on this specific benchmark.
Important point, and often misunderstood: truncating does not reduce the price charged. OpenAI charges based on the model queried, not the size of the vector returned, since the actual cost is the calculation performed on the input text. Requesting 1536 dimensions from text-embedding-3-large therefore costs the same price as its 3072 native dimensions (source: OpenAI, official blog, January 25, 2024); only the storage and search speed change.
This is precisely the default choice of Aurabase, verified in config/mod.rs: the model configured by default is text-embedding-3-large, but the output dimension configured by default is 1536, not 3072. The service therefore pays for the representation of the wide model, truncated to remain on an indexable vector column in native HNSW, without the halfvec cast required to 3072.
Which dimension to choose according to your use case
1536 dimensions remains the reasonable starting point for the majority of RAG or semantic search projects: the MTEB score of text-embedding-3-small (62.3%) remains close to that of the large model, storage remains light, and the classic vector type of pgvector indexes in HNSW without any particular configuration.
3072 dimensions is justified when the corpus is ambiguous or technical, where the quality gap between 62.3% and 64.6% translates into visibly better search results on your own queries, not on the general OpenAI benchmark. Several team feedbacks recorded in the OpenAI developer community discussions point in this direction: the gain of 3072 dimensions is measured on a case-by-case basis, it cannot be assumed.
768 dimensions is especially suitable when volume takes precedence over nuance: a large corpus where the storage or calculation budget is the real constraint, or the use of a legacy embedding model already in 768 native dimensions.
Do not set the dimension until you have measured search quality on a representative sample of your own corpus, not just the general MTEB score published by OpenAI. The MTEB averages dozens of heterogeneous tasks; your RAG corpus is just one.
This choice of dimension is part of a larger RAG stack, embeddings, HNSW index, hybrid search, which our page Native AIdocuments.
What we get asked most often
The right choice is not the greatest, it is the best measured
768, 1536 and 3072 dimensions are not divided on a single axis. 3072 gains in MTEB score published by OpenAI, but leaves pgvector's native HNSW support and pays the price of the large model, whatever the dimension ultimately requested. 1536 remains the most common default balance, including at Aurabase. 768 serves cases where volume takes precedence over nuance.
OpenAI's dimensions parameter changes the question to ask: it is no longer "which model to choose", but "which truncation to accept, for what gain measured on my corpus". Before finalizing a choice in production, test the search quality on a real sample, not just on a general benchmark. Our guide RAG pipeline with pgvector details the complete setup, from ingestion to hybrid search.