Aurabase's native vector search (RAG, pgvector, embeddings) relies on this same indexing mechanism, described in detail on the Native AI on Postgrespage. This guide assumes a Postgres table with pgvector already installed, a column of type vector, and at least a few thousand rows. Below, a simple sequential scan is often faster than an approximate index.
The essentials
- HNSW does not require any training phase, unlike IVFFlat: the index is built over insertions, available in pgvector since version 0.5.0.
- Two parameters set the quality of the index at construction:
m(connections per node, default 16) andef_construction(search width at construction, default 64). - A third parameter,
hnsw.ef_search(pgvector default: 40), is adjusted for each request, without rebuilding the index, to arbitrate recall and latency. - pgvector caps HNSW indexing of type
vectorto 2000 dimensions. Beyond that (an embedding with 3072 dimensions, for example), a cast tohalfvecis necessary to index. - pgvector 0.8.6 is the version embedded in the Aurabase Postgres tenant image, verified directly in the Dockerfile on August 24, 2026.
What is an HNSW index in pgvector?
HNSW stands for Hierarchical Navigable Small World. It is a graph index: each vector becomes a node connected to its closest neighbors, organized in several superimposed layers. A search starts at the top of the graph, on the sparsest layer, then goes down layer by layer to the most relevant neighbors. The search time thus becomes almost logarithmic, not linear over the number of lines.
IVFFlat, the other index of pgvector, works differently: it divides the vector space into lists determined by a training pass on an existing sample, before being able to index anything. HNSW does not have this constraint, each insertion directly enriches the graph, which makes it simpler to operate on a table that grows continuously. On the other hand, an HNSW index consumes more memory and takes more time to build than an equivalent IVFFlat on the same volume.
pgvector introduces HNSW support in version 0.5.0. Later versions add useful capabilities for this guide: the halfvec type (0.7.0) to index beyond 2000 dimensions, and the hnsw.iterative_scan parameter (0.8.0) to improve recall on filtered queries. If you compare pgvector to a dedicated vector base before deciding, our comparison pgvector vs Pinecone, Weaviate and Qdrant details the trade-offs.
Check your version of pgvector before creating the index
Confirm the version of pgvector installed first. An extension that is too old causes some features in this guide to silently fail, particularly halfvec and hnsw.iterative_scan.
HNSW has existed since pgvector 0.5.0. The halfvectype, necessary to index embeddings beyond 2000 dimensions, requires at least version 0.7.0. The hnsw.iterative_scan parameter requests version 0.8.0.
On Aurabase projects, the question does not arise: the Postgres image embeds pgvector 0.8.6, both on the shared Postgres cluster (docker/Postgres.Dockerfile, built directly on pgvector/pgvector:0.8.6-pg16-bookworm) and on the Postgres 16 CNPG instances dedicated per project (docker/Postgres.CNPG.Dockerfile, which inherits pgvector 0.8.6 from the official CloudNativePG image). Verified in both Dockerfiles on August 24, 2026.
Choose the right type of column according to the size of your embeddings
The column type depends on the size of your embeddings, not just the model that generates them. pgvector stores a classic vector in the vectortype, with a cap of 16,000 dimensions in storage. But HNSW indexing on this type is limited to 2000 dimensions: beyond that, CREATE INDEX fails.
Common embedding models often exceed this threshold: text-embedding-3-large from OpenAI or gemini-embedding-2 from Google natively produce up to 3072 dimensions. To index these vectors with HNSW, cast the column to halfvec (storage precision halved), which pushes the indexing limit well beyond 2000 dimensions.
| Dimensions | Column | HNSW on vector | Requires casting |
|---|---|---|---|
| 768 | embedding_768 | Yes | No |
| 1536 | embedding_1536 | Yes | No |
| 3072 | embedding_3072 | No (> 2000 dims) | Yes, cast::halfvec(3072) |
The Aurabase RAG engine illustrates this compromise in production: three classes of dimensions supported (768, 1536, 3072), stored in three distinct columns of the same embeddingstable. Columns 768 and 1536 are indexed directly in HNSW on the vectortype. Column 3072 is indexed via a ::halfvec(3072)cast, precisely to circumvent the 2000 dimension cap.
For details of ingestion (chunking, call to the embedding provider, insertion), see the RAG pipeline tutorial on pgvector.
Create the index with the m and ef_construction parameters
The minimal syntax is sufficient for a first index, with the default values of pgvector.
pgvector then applies m = 16 and ef_construction = 64. To adjust these values explicitly, use the WITH clause:
Before building an HNSW index on a large table, temporarily increase maintenance_work_mem for the session: this is, according to the pgvector documentation itself, the most direct lever to reduce construction time.
What does the parameter m change?
m sets the maximum number of connections that each node in the graph maintains per layer. A higher value densifies the graph: recall increases, but memory consumed and construction time also increase, approximately linearly. The default (16) is suitable for most cases. Going up to 24 or 32 is especially justified on large embeddings, where the distinction between close and distant neighbors becomes finer.
What does ef_construction change?
ef_construction sets the size of the candidate list explored during index construction, for each inserted node. A higher value improves the quality of the final graph, therefore the potential recall, at the cost of a longer construction time. Unlike m, this parameter has no cost at query time: it is a one-time investment, paid only once when the index is created.
Partial indexes for several dimension classes in the same table
When a table stores multiple vector columns (one per dimension class, as Aurabase does), index each column separately with an WHERE colonne IS NOT NULLclause. This partial index avoids indexing empty lines for classes not used by a given line, which reduces the size of the index and speeds up its construction without costing anything in recall.
The choice of operator class (vector_cosine_ops, vector_l2_ops or vector_ip_ops) must correspond to the metric on which the embedding model was trained. Most recent text embedding models are trained for cosine similarity: vector_cosine_ops (or halfvec_cosine_ops on a cast column) is therefore the safest default choice.
Set ef_search at query time
ef_search is set on each query, not when the index is built. It sets the size of the list of candidates explored during the search: the higher it is, the better the recall, at the cost of longer latency. pgvector sets its default value to 40.
40 is rarely enough as soon as a query combines vector search with a WHERE filter applied after scanning the index (on a namespace, a tenant, or any other metadata criterion). The HNSW scan brings back ef_search raw candidates, then the filter discards part of them. If too few candidates survive, the final LIMIT ends up underfilled.
The Aurabase RAG engine therefore expands ef_search dynamically according to the requested top_k, instead of keeping the fixed value of 40: ef = max(top_k × 4, 64). A search for the 5 closest results uses ef_search = 64; a search for the top 50 uses ef_search = 200. This formula remains adjustable per environment variable for deployments that need another recall/latency tradeoff.
pgvector 0.8 adds a second lever for this same problem: hnsw.iterative_scan. In strict_order or relaxed_ordermode, the search gradually broadens its search until it gathers enough results after filtering, rather than stopping on a fixed list of candidates. Aurabase activates it by default in strict_order, but protects the call in a savepoint. On a version of pgvector before 0.8, where this parameter does not exist, the query continues in degraded mode rather than failing.
Build the complete RAG pipeline
This HNSW index is just one piece of the complete RAG pipeline: chunking, embedding generation, ingestion, then search. Our step-by-step tutorial builds this pipeline end-to-end on pgvector, from the first insertion to the similarity query. The technical documentation also details all of Aurabase's native AI capabilities built on Postgres.