Without an index, pgvector does a sequential scan and returns exact nearest neighbours: perfect recall, and a query time that grows with every row you insert. Add an index and recall stops being guaranteed, because both major options in the HNSW vs IVF-Flat choice are approximate by design. The question is not which one is more accurate in the abstract. It’s which approximation you can afford at your query volume and your rebuild schedule.
HNSW builds a layered graph with no training step, needs more memory and a slower build, and recovers recall at query time by raising ef_search, with cost that grows sublinearly. IVF-Flat clusters vectors into lists during a training step that requires data already in the table, builds faster on less memory, and recovers recall by raising probes, with cost that grows linearly. For production recall targets, HNSW is the default; IVF-Flat is the fallback when build time or memory is the binding constraint.
What actually separates HNSW vs IVF-Flat at query time?
IVF-Flat partitions the vector space into clusters, commonly called lists, during index build. A query first finds the nearest cluster centroids, then does an exact scan inside however many clusters probes tells it to check, following the same lists-and-probes tuning pattern documented across pgvector deployments. Miss the right cluster and the true neighbour is gone from the result set no matter how good the in-cluster scan is, which is why recall rises roughly linearly with probes and query time rises right alongside it.
HNSW, described in the original Malkov and Yashunin paper, never partitions anything. It builds a multi-layer graph and walks it: a sparse top layer for long jumps, denser layers below for local refinement, similar in spirit to a skip list built over similarity instead of order. That structure is what gives it logarithmic-ish search behaviour instead of the linear scan-more-clusters pattern IVF-Flat depends on, and it’s why partitioning-based scaling and graph-based scaling age so differently as a dataset grows.
| Property | HNSW | IVF-Flat |
|---|---|---|
| Build step | None required; can build on an empty table. | Training step; needs representative data first. |
| Query-time knob | ef_search, tuned per query. |
probes, tuned per query. |
| Recall vs latency | Sublinear; recall rises faster than cost. | Roughly linear; cost tracks recall directly. |
| Build cost | Slower, more memory per pgvector’s own docs. | Faster, lower memory footprint. |
1-- HNSW: no training step, tunable at build and query time 2CREATE INDEX ON items USING hnsw (embedding vector_cosine_ops) 3 WITH (m = 16, ef_construction = 64); 4SET hnsw.ef_search = 100; 5 6-- IVF-Flat: build only after loading representative data 7CREATE INDEX ON items USING ivfflat (embedding vector_cosine_ops) 8 WITH (lists = 100); 9SET ivfflat.probes = 10; -- pgvector suggests sqrt(lists) as a start
HNSW vs IVF-Flat: which one should you actually pick?
Reach for HNSW first. The pgvector project itself describes it as having the better speed-recall tradeoff of the two, at the cost of slower builds and more memory, and it needs no training step so it can be built the moment the table exists. IVF-Flat earns its place when the build itself is the bottleneck: it is faster to construct and lighter on memory, which matters on a large table you rebuild often, but it must be built after loading data, since its clusters are only as good as the distribution they were trained on. Retrain it after any major shift in the data, or the cluster boundaries stop matching reality.
Either index sits underneath the same job: surfacing the right neighbours for whatever consumes them next, commonly a RAG pipeline’s retrieval stage. The index choice does not change what gets built on top of it, only how reliably that stack gets the right rows to work with.
- Without any index, pgvector performs an exact sequential scan with perfect recall; both HNSW and IVF-Flat trade some of that recall for speed.
- HNSW needs no training step, recovers recall through ef_search, and scales sublinearly, at the cost of slower builds and higher memory use.
- IVF-Flat requires a training step on existing data, recovers recall through probes, and scales roughly linearly with query cost.
- pgvector’s own documentation recommends HNSW for its query performance and reserves IVF-Flat for build-time or memory-constrained deployments.
- An IVF-Flat index trained on one data distribution degrades as that distribution shifts, and needs rebuilding to stay accurate.
Conclusion
Default to HNSW unless a specific constraint, usually build time or memory, pushes you toward IVF-Flat. Either way, measure recall against exact search on a validation set before trusting the index in production. For what typically sits downstream of this index, see RAG architecture: a complete guide.
Frequently Asked Questions
What is the main difference between HNSW and IVF-Flat?
HNSW builds a multi-layer graph and searches it by walking from sparse upper layers to denser lower ones, with no training step required. IVF-Flat clusters vectors into lists during a training step that needs existing data, then searches by checking the nearest clusters at query time. Their recall-tuning knobs also differ: ef_search for HNSW, probes for IVF-Flat.
Which index has better recall for production use?
pgvector’s own documentation describes HNSW as having the better speed-recall tradeoff between the two, at the cost of slower index builds and more memory. Most production deployments default to HNSW unless build time or memory is a hard constraint that pushes them toward IVF-Flat instead.
Why does IVF-Flat need data before you can build the index?
IVF-Flat runs a clustering step at build time to decide where the list centroids sit, and that clustering needs representative data to produce useful clusters. Building it on an empty or unrepresentative table produces poorly placed centroids, which shows up later as unexpectedly low recall.
How do I recover recall if it drops after adding an index?
Raise the query-time tuning parameter for whichever index you’re using: ef_search for HNSW or probes for IVF-Flat. Both trade query latency for recall, but HNSW’s cost grows sublinearly as you raise ef_search while IVF-Flat’s cost grows roughly linearly with probes.
Does an IVF-Flat index need to be rebuilt over time?
Yes, if the underlying data distribution shifts meaningfully. The cluster centroids are fixed at build time based on the data present then, so a table that grows or changes significantly afterward can develop clusters that no longer match its real distribution, degrading recall until the index is rebuilt.