What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose HNSW when query speed and recall matter more than index build time and memory; choose IVFFlat when you need a lighter, quicker-to-build index and can train and tune its lists. That is pgvector’s project-level guidance, not a guarantee for every workload. Both indexes are approximate, so benchmark against exact search using your data, filters, and latency goals.

What changes when you choose one index over the other?

pgvector performs exact nearest-neighbor search by default. HNSW and IVFFlat are approximate indexes: they can make searches faster while sacrificing some recall, meaning they may not return every nearest neighbor that exact search would find.

Decision factor HNSW IVFFlat
Speed-recall tradeoff The pgvector project describes its query performance as better than IVFFlat in the speed-recall tradeoff. The project describes its query performance as lower in that tradeoff.
Build and memory Slower to build and uses more memory. It does not require a training step, so it can be created before data is loaded. Faster to build and uses less memory. Build it after loading data so its lists can be trained from the table.
Main controls m (default 16), ef_construction (default 64), and query-time ef_search (default 40). lists and query-time probes (default 1).
Filtered queries Filtering occurs after the approximate scan. Iterative scans can search further within configured limits. Filtering also occurs after the scan. Iterative scans can continue up to ivfflat.max_probes.

These are qualitative comparisons and documented defaults from the pgvector README, not universal performance measurements. Results depend on data, hardware, query shape, filters, and configuration.

When should you choose HNSW?

Prefer HNSW when search quality and latency are priorities

HNSW builds a multilayer graph and is the natural starting point when the workload benefits from stronger query performance in the speed-recall tradeoff and can accommodate higher memory use and slower index creation. The project summarizes the tradeoff this way: “It has better query performance than IVFFlat (in terms of speed-recall tradeoff), but has slower build times and uses more memory.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Understand its controls

  • m controls the graph’s connections; the documented default is 16.
  • ef_construction controls the construction search effort. Raising it can improve recall, but increases build time and reduces insert speed. Its documented default is 64.
  • ef_search controls query-time search effort. Raising it can improve recall at the cost of query speed. Its documented default is 40.

The README advises starting with defaults unless recall is low. HNSW builds faster when the graph fits in maintenance_work_mem, but do not raise that setting so far that it exhausts server memory.

When should you choose IVFFlat?

Prefer IVFFlat when build cost and memory are important

IVFFlat clusters vectors into lists, then searches a subset near the query vector. It builds faster and uses less memory than HNSW, but requires training from table data. Load data before creating the index; an index made with too little or unrepresentative data may not serve the eventual workload well.

Set lists and probes as starting points, not guarantees

The README suggests these initial list-count heuristics:

  • Up to 1 million rows: start with lists = rows / 1000.
  • Above 1 million rows: start with lists = sqrt(rows).

Start with probes around sqrt(lists). Increasing probes can improve recall while slowing queries. If probes equals the total number of lists, the search is exact and the planner will not use the IVFFlat index, so that setting defeats the index’s intended approximate-search role.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do you create the indexes correctly?

These examples use cosine distance; use the operator class that matches the distance operator in your queries. pgvector also supports other distance operators, including L2 and inner product. Check the README for requirements that apply to your vector type, dimensions, and installed extension version.

CREATE INDEX ON items USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON items USING ivfflat (embedding vector_cosine_ops)
WITH (lists = 100);

The IVFFlat value of 100 is an illustrative SQL setting, not a recommended value for every table. Choose lists from your data size as a starting point, then measure.

How do filters affect approximate results?

With either index, pgvector applies ordinary filtering after the approximate index scan. A selective condition can therefore leave fewer rows than the requested result count, even if the query asks for more.

Starting with pgvector 0.8.0, iterative index scans can keep scanning until enough qualifying rows are found or the configured limit is reached. Strict ordering preserves exact distance order; relaxed ordering can improve recall while allowing slight deviations in distance order. Confirm that the installed version supports the setting you intend to use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • For a small number of distinct filter values, consider a partial index.
  • For many distinct values, consider partitioning.
  • In multitenant systems, a shared approximate index can let one tenant’s vectors affect another tenant’s recall and speed. The project suggests list partitioning or separate tables for tenant isolation.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How should you compare them on your workload?

Measure recall against exact results, then inspect actual query plans and index-build progress. There is no universal latency or recall figure that establishes a winner for all deployments.

  1. Establish an exact baseline. In a transaction, disable index scans as shown in the pgvector README and run representative queries to capture exact nearest-neighbor results.
  2. Test approximate configurations. Compare HNSW and IVFFlat using the same representative data, distance operator, result count, filters, and query mix. Record recall alongside latency; a fast query that misses too many relevant neighbors may not meet the application’s needs.
  3. Tune the relevant control. For HNSW, test ef_search for query-time recall and speed, and adjust ef_construction if the build/insert tradeoff warrants it. For IVFFlat, test list count and probes.
  4. Inspect execution. Use EXPLAIN (ANALYZE, BUFFERS) to understand query behavior, including whether the planner uses the index and what the scan costs.
  5. Track index creation. Monitor builds with pg_stat_progress_create_index.

The documented guidance and settings are in the official pgvector README. The README includes features introduced in 0.8.0, and its current page references a v0.8.6 release; verify the extension version actually deployed before relying on release-sensitive features or configuration.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.