To improve pgvector search, measure approximate results against PostgreSQL’s exact nearest-neighbor results, then tune the index and search effort against your real queries and filters. HNSW and IVFFlat can make searches faster, but both trade some recall for speed; no single setting is best for every dataset. The pgvector project says its default, unindexed exact search provides perfect recall, making it a useful baseline (official pgvector README, accessed October 4, 2026).
Build an exact-search baseline first
Before changing an index or query setting, record how representative searches perform without an approximate index. Keep the query vectors, filters, requested result count, data snapshot, and relevant workload conditions consistent when comparing configurations. Record returned row identities as well as execution time: latency alone cannot show whether a faster query returned the same nearest neighbors.
- Measure a representative query. Run it with the filters and result count used by your application. Use
EXPLAIN (ANALYZE, BUFFERS)to inspect the actual plan, execution information, and buffer activity. - Run the exact comparison. In a transaction, use
SET LOCAL enable_indexscan = off;before running the query. The pgvector README documents this as a way to compare approximate results with exact search. - Save the baseline. Record the exact result identities and the query’s execution details. Use these results to assess recall and speed after each controlled change.
Compare like with like. Changing the data, filter, query vector, or result count between runs makes it harder to tell whether a tuning change helped.
Choose HNSW or IVFFlat for your constraints
Both are approximate-nearest-neighbor indexes, but they make different tradeoffs. The pgvector README describes their relative behavior qualitatively; it does not establish a universal winner or performance figures for a particular application.
Recommended Free Tools
#1 Best Overall
| Index | Documented tradeoff | When to consider it |
|---|---|---|
| HNSW | Generally better query performance in the speed/recall tradeoff, but slower index builds and higher memory use than IVFFlat. It has no IVFFlat-style training step and can be created before the table contains data. | When its query performance and recall tradeoff suits the workload and the memory and build costs are acceptable. |
| IVFFlat | Faster builds and lower memory use, but lower query performance in the speed/recall tradeoff than HNSW. It divides vectors into lists and searches a subset near the query. | When faster builds and lower memory use matter, and there is data available before index creation. |
Compare both on the things your deployment needs: recall at the required result count, latency, memory footprint, build time, insertion or refresh patterns, and behavior under production-like filters.
Tune IVFFlat lists and probes
IVFFlat needs data before index creation. The pgvector README gives starting heuristics for the number of lists, not guaranteed optimal settings for every dataset or query mix:
Rank #2
- For up to one million rows, start around
rows / 1000lists. - Above one million rows, start around the square root of the row count.
- Start
ivfflat.probesaround the square root of the list count, then benchmark.
Increasing probes searches more lists and improves recall at the cost of speed. The README states that setting probes equal to the number of lists reaches exact nearest-neighbor search; at that point PostgreSQL will not use the IVFFlat index. Test probe values with the same queries and exact baseline rather than assuming the heuristic is optimal.
Tune HNSW search effort and iterative scans
The documented default for hnsw.ef_search is 40. A limited candidate list, dead tuples, and filters can contribute to getting fewer results than requested. If comparison shows poor recall or queries return too few results, raise search effort and measure the latency cost.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #3
Starting with pgvector 0.8.0, iterative index scans can continue scanning until enough results are found or a scan limit is reached. They support two ordering modes:
- Strict ordering: preserves exact distance order.
- Relaxed ordering: may improve recall while allowing results to be slightly out of order. A materialized CTE can restore strict ordering afterward. For PostgreSQL 17 and later, the README’s example requires
+ 0in the outer ordering expression.
For HNSW iterative scans, documented controls include hnsw.max_scan_tuples, which defaults to 20,000, and hnsw.scan_mem_multiplier, which defaults to 1. For IVFFlat, ivfflat.max_probes caps probes during iterative scans. Higher scan limits can cost time or memory, so evaluate them alongside recall and latency.
Understand why WHERE filters can reduce results
With an approximate index, pgvector applies filters after scanning the index. A selective WHERE condition can therefore leave fewer qualifying rows than the requested result count. For illustration, the pgvector README says that a condition matching 10% of rows would yield an average of four matching rows with HNSW’s default hnsw.ef_search of 40. This is an example, not a general guarantee.
Choose a filtering strategy based on how selective the condition is and how many values it has:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Low percentage of rows matches: a conventional index on the filter column can enable fast exact nearest-neighbor search in many cases.
- Several filter columns: consider a multicolumn index.
- Only a few filter values: consider a partial approximate index.
- Many distinct values: consider partitioning.
- Approximate results need more matches: iterative scans can continue scanning for enough qualifying results, subject to their scan limits.
Account for tenant isolation
In a multitenant application, a shared approximate index can let one tenant’s vectors affect another tenant’s recall and speed. The pgvector README suggests list partitioning or separate tables when isolation is needed. Test filtered, tenant-specific queries separately from unfiltered searches; a good aggregate result can hide poor behavior for a particular tenant.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Consider storage and index-maintenance tradeoffs
For a smaller working set, pgvector describes halfvec as a lower-precision storage option. Binary quantization can make indexes smaller and speed builds at scale; reranking binary-search candidates with the original vectors is a documented way to improve recall. These options affect precision and ranking, so compare result quality as well as performance on representative queries.
For large initial loads, the README recommends bulk loading with COPY and creating indexes afterward. Increasing parallel maintenance workers can speed index creation. In production, CREATE INDEX CONCURRENTLY avoids blocking writes. HNSW vacuuming can take a while; the project’s documented suggestion is to reindex concurrently before vacuuming.
Quick Recap
Use a controlled tuning loop
- Capture representative query vectors, filters, requested result counts, and exact-search results.
- Choose HNSW or IVFFlat based on the measured needs for recall, latency, memory, build time, and data changes.
- Change one setting at a time: IVFFlat lists or probes, HNSW search effort, or iterative-scan limits.
- Benchmark filtered searches independently. Depending on selectivity and value cardinality, test a conventional filter index, multicolumn index, partial index, partitioning, or iterative scans.
- Compare approximate result identities with exact results and inspect plans using
EXPLAIN (ANALYZE, BUFFERS). For ongoing query behavior, PostgreSQL tools such aspg_stat_statementsor PgHero can help monitor workloads. - Keep tested settings tied to the data size and workload they were measured against. Revisit them when volume, filters, concurrency, or latency requirements change.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

