HNSW in Postgres with pgvector: how to evaluate retrieval quality
Evaluate Postgres HNSW vector search with exact-search recall, latency, filters, memory, index build cost, ef_search, reranking, and downstream RAG quality.
HNSW in Postgres trades exact nearest-neighbor results for faster approximate search. Evaluate it by comparing HNSW results with exact search on representative queries, then measure latency, recall, filtering behavior, memory, build time, and downstream RAG quality.
The maintained Postgres extension is
pgvector; older pg_embedding material
may no longer describe the stack you deploy. Pin the extension version.
Create an exact baseline
For each query, run exact nearest-neighbor search with the same distance metric, filters, and candidate limit. Treat its top-k IDs as the retrieval baseline. Then run HNSW and calculate recall@k and overlap.
Use production-like embeddings, corpus size, filters, and query distribution. Random vectors do not reveal semantic or tenant-filter behavior.
Tune the speed-recall tradeoff
pgvector exposes m and ef_construction for index construction and ef_search
for queries. Higher search effort generally improves recall at a latency cost.
Sweep settings and plot p50/p95 latency against recall.
Use the operator class that matches cosine, inner-product, or L2 semantics. A metric mismatch invalidates the comparison.
Evaluate filters
Approximate index filtering can return too few rows because filtering is applied to scanned candidates. Test tenant, category, time, and permission filters at real selectivity. Consider iterative scans, partial indexes, partitioning, or exact search for highly selective cases according to pgvector's documentation.
Never evaluate unfiltered recall and assume tenant-filtered queries behave the same.
Measure operations
Record index build time, index size, write latency, vacuum behavior, memory, and
replica impact. Test index creation and upgrades using production-safe procedures.
Monitor with Postgres query statistics and EXPLAIN (ANALYZE, BUFFERS).
Evaluate downstream RAG
Retrieval overlap does not prove answer quality. Run the same RAG dataset with exact and HNSW retrieval, then compare groundedness, citations, latency, and cost. Trace query embedding, SQL settings, retrieved IDs, reranking, and final generation.
Currai can store these retrieval spans under the answer trace, allowing a quality
regression to be segmented by index version and ef_search. That connects
database tuning to the user-visible outcome.
