Sep 28, 2026

HNSW in Postgres with pgvector: how to evaluate retrieval quality

Evaluate Postgres HNSW vector search with exact-search recall, latency, filters, memory, index build cost, ef_search, reranking, and downstream RAG quality.

TUTORIAL10 min readThe Currai team / Engineering

HNSW in Postgres trades exact nearest-neighbor results for faster approximate search. Evaluate it by comparing HNSW results with exact search on representative queries, then measure latency, recall, filtering behavior, memory, build time, and downstream RAG quality.

The maintained Postgres extension is pgvector; older pg_embedding material may no longer describe the stack you deploy. Pin the extension version.

Create an exact baseline

For each query, run exact nearest-neighbor search with the same distance metric, filters, and candidate limit. Treat its top-k IDs as the retrieval baseline. Then run HNSW and calculate recall@k and overlap.

Use production-like embeddings, corpus size, filters, and query distribution. Random vectors do not reveal semantic or tenant-filter behavior.

Tune the speed-recall tradeoff

pgvector exposes m and ef_construction for index construction and ef_search for queries. Higher search effort generally improves recall at a latency cost. Sweep settings and plot p50/p95 latency against recall.

BEGIN;
SET LOCAL hnsw.ef_search = 100;
SELECT id
FROM documents
ORDER BY embedding <=> $1
LIMIT 10;
COMMIT;

Use the operator class that matches cosine, inner-product, or L2 semantics. A metric mismatch invalidates the comparison.

Evaluate filters

Approximate index filtering can return too few rows because filtering is applied to scanned candidates. Test tenant, category, time, and permission filters at real selectivity. Consider iterative scans, partial indexes, partitioning, or exact search for highly selective cases according to pgvector's documentation.

Never evaluate unfiltered recall and assume tenant-filtered queries behave the same.

Measure operations

Record index build time, index size, write latency, vacuum behavior, memory, and replica impact. Test index creation and upgrades using production-safe procedures. Monitor with Postgres query statistics and EXPLAIN (ANALYZE, BUFFERS).

Evaluate downstream RAG

Retrieval overlap does not prove answer quality. Run the same RAG dataset with exact and HNSW retrieval, then compare groundedness, citations, latency, and cost. Trace query embedding, SQL settings, retrieved IDs, reranking, and final generation.

Currai can store these retrieval spans under the answer trace, allowing a quality regression to be segmented by index version and ef_search. That connects database tuning to the user-visible outcome.

03

Keep going with nearby topics from the Currai blog.