HNSW and IVF are both approximate-nearest-neighbor index families; neither is universally faster or more accurate. HNSW navigates a graph and is a strong first benchmark when the index fits in RAM. IVF searches selected partitions and offers choices from full-precision IVF-Flat to more compact, lower-recall IVF-PQ. Choose by measuring recall at your latency target while tracking memory and build costs.
How HNSW and IVF find neighbors
HNSW: navigate a graph
Hierarchical Navigable Small World (HNSW) connects vectors in a layered graph. A query follows links toward likely neighbors rather than comparing against every vector. The links add memory overhead, but can make approximate search fast. FAISS describes HNSW’s M as the number of links per vector and documents a range of 4–64; higher values use more RAM. These are FAISS-specific configuration notes, not universal limits. See FAISS’s index-selection guidance. The method’s foundational paper is listed in the HNSW paper record.
IVF: probe selected partitions
An inverted-file index (IVF) assigns vectors to coarse clusters, then searches selected clusters for a query. IVF is a family, not a single memory or accuracy profile: IVF-Flat retains full-precision vectors, while IVF-SQ and IVF-PQ store compressed representations. More compression can reduce storage and memory bandwidth, but can also reduce recall and require additional tuning or reranking. FAISS discusses these variants in its index-selection guidance; NVIDIA cuVS outlines their trade-offs in its indexing guide.
They can be combined
HNSW and IVF are not always mutually exclusive. FAISS’s large-scale indexing guide compares IVF approaches that use HNSW as a coarse quantizer alongside other quantizers.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute#1 Best Overall
Trade-offs that matter in practice
| Decision axis | HNSW | IVF | What to measure |
|---|---|---|---|
| Search structure | Layered graph traversal. | Coarse clusters and inverted lists; search probes selected partitions. | Candidate work and end-to-end latency on your workload. |
| Memory | Vector storage plus graph links; link configuration affects overhead. | IVF-Flat stores full vectors plus partition metadata; quantized IVF stores compact representations. | Total resident memory for the actual implementation and data type. |
| Recall and latency tuning | In FAISS, efSearch adjusts search effort; more effort can improve quality at a speed cost. |
In FAISS, nprobe controls how many partitions are searched; more probes increase work and can improve recall. |
Recall@k at the required latency, not latency in isolation. |
| Compression | Graph indexing alone does not provide IVF-PQ-style compression; compression is a separate design choice. | IVF-SQ and IVF-PQ trade representation size and bandwidth against accuracy; PQ compresses more aggressively and needs tuning. | Recall loss, bytes per vector, and any refinement or reranking cost. |
| Build and training | FAISS says HNSW does not require training, though graph construction can be expensive. | IVF requires clustering; IVF-PQ also trains codebooks. | Build time, training resources, and rebuild frequency. |
| Operational fit | A candidate when the index fits in RAM and high-quality CPU search is desired. | A candidate when partitioning or a lower footprint matters, especially with IVF-PQ when index size is the bottleneck. | Update and filter behavior, available memory, workload mix, and implementation-specific details. |
FAISS gives a useful HNSW memory model for its stated representation: (d * 4 + M * 2 * 4) bytes per vector, where d is vector dimension and M controls graph links. Treat that as an illustration of graph overhead, not a universal estimate for every library or database. Check the FAISS guidance and calculate memory for your exact index and data type.
Which index should you benchmark first?
Start with HNSW when RAM is available
If the index comfortably fits in memory and CPU query quality is the priority, benchmark HNSW first. Tune efSearch against your recall and latency goals. FAISS recommends HNSW for small datasets or systems with ample RAM, and NVIDIA cuVS likewise recommends it for high-quality CPU search when the index fits in memory. Those are implementation-specific starting points, not performance guarantees. See FAISS and cuVS.
Rank #2
Choose IVF-Flat when memory is tight but precision matters
IVF-Flat retains full-precision vectors while limiting query work to selected partitions. It is worth benchmarking when graph overhead is a concern but compression is undesirable. Use nprobe to tune the amount of the partitioned index examined, then check recall at the resulting latency.
Consider quantized IVF when footprint is the constraint
IVF-SQ and IVF-PQ reduce the stored representation compared with full-precision vectors, with accuracy trade-offs. cuVS characterizes IVF-SQ as a smaller recall trade-off and IVF-PQ as a choice when index size is the main bottleneck and stronger compression is worth additional tuning or reranking. Measure the recall impact and any reranking work on your own queries before committing.
Recommended Free Tools
Keep an exact baseline if exact neighbors are required
Approximate indexes trade search work for speed or footprint; they do not guarantee exact results. FAISS says its Flat indexes are the only indexes in its guidance that guarantee exact results and recommends them as a baseline for evaluating approximate indexes. See FAISS’s guidance.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to compare them fairly
- Fix the workload. Use the same dataset, vector dimensions, distance metric, representative query set, filtering, hardware, and concurrency for each candidate.
- Set the quality target. Choose the relevant
kand required recall@k before tuning. Compare candidates at matched recall, rather than declaring a winner from latency alone. - Tune the search controls. For FAISS HNSW, vary
efSearch; for FAISS IVF, varynprobe. Measure the recall and latency at each setting. - Record serving and lifecycle costs. Track p50, p95, and p99 latency, throughput under expected concurrency, peak memory, build or training time, and update behavior.
- Use the result for your deployment, not as a universal ranking. Defaults and behavior vary across implementations and versions, including how filtering, updates, and memory tiers work. Verify the official documentation for the library and version you will run.
Published benchmark numbers are meaningful only with their setup attached. For example, FAISS’s large-scale indexing guide reports one operating point at nprobe=128 and quantizer_efSearch=32: recall@1 of 0.6786 and 0.05387 ms/query for that experiment. The guide says its experiments used a normalized 2.2 GHz Xeon E5-2698, 80-core platform and ran with 32 cores. This is not a general IVF speed or recall figure. See the FAISS experiment details.
Rank #4
What the evidence does—and does not—establish
The official FAISS and NVIDIA cuVS guidance provides useful starting points for choosing and tuning index configurations, but it does not establish a workload-independent HNSW-versus-IVF winner for latency or recall. Results depend on dataset, dimension, metric, index variant, tuning, hardware, and target recall. The comparison here concerns algorithm families; an individual database may change defaults, update behavior, filtering strategies, supported metrics, or where vectors and indexes reside. Confirm those details for your product and version before selecting parameters.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →

