Aurabase Logo
aurabasedocs
docs›Guides›RAG & pgvector

RAG & Vector Search Guide

Build a lightning-fast semantic search and comprehensive Retrieval-Augmented Generation (RAG) system combining PostgreSQL pgvector and Aurabase's unified AI Gateway.

12 min read·Level advanced·Revised on Aug 19, 2026
This English text was generated automatically from the French original and has not been reviewed yet.
#
Architecture

RAG on Aurabase in 3 steps

RAG allows your AI models to respond accurately on your own private data without retraining:

  • Vectorization: Your textual documents are transformed into digital vectors using aura.ai.embed().
  • Storage & Indexing: Vectors are stored in PostgreSQL with the native extension pgvector.
  • Augmentation & Answer: The question is compared to the stored vectors, and the relevant snippets are provided to the LLM via aura.ai.chat().
#
Implementation

Full RAG Pipeline Code

src/lib/index-doc.tstypescript
import { aura } from '@/lib/aurabase'

export async function indexDocument(content: string, metadata: Record<string, any>) {
  // Calculation of vector embedding via the AI Gateway Aurabase
  const { vectors } = await aura.ai.embed({
    provider: 'openai',
    model: 'text-embedding-3-small',
    input: [content],
  })

  // Saving in PostgreSQL table
  const { data, error } = await aura.from('documents').insert({
    content,
    metadata,
    embedding: vectors[0],
  })
  return { data, error }
}
#
Performance

HNSW indexing for sub-10ms queries

For databases containing hundreds of thousands of vectors, HNSW (Hierarchical Navigable Small World) indexing provides significantly better search performance than ivfflat:

HNSW Index (Recommended)

Zero pre-training time, recall > 99% and constant latency even with continuous insertions of new documents.

Cosine Distance (<=>)

Use the <=> operator to calculate the normalized cosine distance between the embedding vectors.

#
Optimization

Semantic Cache & Token Savings

Aurabase's AI Gateway integrates an automatic semantic cache: if two users ask a similar question (cosine similarity greater than 0.95), the answer is served from the in-memory cache in less than 2 ms, saving 100% of the LLM call cost.

Last updated · Aug 19, 2026