Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Start by running EXPLAIN (ANALYZE, BUFFERS) on the slow query with representative parameters. The plan shows whether PostgreSQL uses the intended vector index, how many rows it examines, and where time and buffer activity accumulate. Then check query shape, approximate-index settings, filters, and maintenance—while measuring result quality as well as speed.
1. Capture the real query plan
Use the slow query itself, with realistic parameter values and representative data volume:
EXPLAIN (ANALYZE, BUFFERS)
SELECT *
FROM items
ORDER BY embedding <-> '[3,1,2]'
LIMIT 5;
Inspect actual elapsed time, buffer activity, row counts, and the plan node used to retrieve nearest neighbors. A sequential scan is not automatically a problem: on a small table it can be faster than using an index. ANALYZE executes the query, so run it where executing that query is appropriate. The pgvector project README recommends examining plans with ANALYZE and BUFFERS.
2. Confirm the query can use a vector index
For the documented indexable pattern, order by a distance operator in ascending order and apply a LIMIT. For example:
#1 Best Overall
SELECT *
FROM items
ORDER BY embedding <-> '[3,1,2]'
LIMIT 5;
A transformed expression such as ORDER BY 1 - (embedding <=> query) DESC does not match that documented form. If the plan still chooses a sequential scan, you can test whether an index plan is possible by temporarily discouraging sequential scans in a transaction:
BEGIN;
SET LOCAL enable_seqscan = off;
EXPLAIN (ANALYZE, BUFFERS)
SELECT *
FROM items
ORDER BY embedding <-> '[3,1,2]'
LIMIT 5;
ROLLBACK;
This is a diagnostic experiment, not a production setting to apply indiscriminately. If forcing the planner’s preference reveals a usable index plan, check the original query form and whether the table is large enough for index use to pay off.
3. Decide whether the bottleneck is exact or approximate search
pgvector performs exact nearest-neighbor search by default. Exact search has perfect recall, but its cost can grow with the dataset. HNSW and IVFFlat indexes make search approximate: they can reduce query cost, but may return different neighbors. Compare them with exact results on a representative sample rather than judging an index by latency alone.
Rank #2
To obtain an exact-search comparison when an approximate index exists, the project README documents locally disabling index scans with SET LOCAL enable_indexscan = off inside a transaction. Compare the returned neighbors and query latency against the indexed query. Track recall or another application-appropriate quality measure alongside elapsed time and buffer reads.
Free tools Windows power users keep installed
One-click scans. No signup required.
When exact search is slow
- For an exact scan without an index, the project suggests increasing
max_parallel_workers_per_gather. Validate the effect on your workload; it is not a universal speed guarantee. - If your vectors are normalized to length 1, test inner product as a distance measure. The project notes it can be faster for normalized unit vectors.
- If the query is selective because of a filter, inspect whether an ordinary index on the filter column can narrow the candidate rows before exact distance ordering.
When approximate search is slow or misses useful results
Tune the selected index’s search breadth or probes, then compare speed and recall against the exact baseline. The index type and settings affect the balance differently; use the next section for the relevant controls.
4. Tune the approximate index you actually use
| Index | What to check | Trade-off and next experiment |
|---|---|---|
| HNSW | hnsw.ef_search; for filtered queries, iterative scanning and its limits |
Increasing search breadth can improve recall while increasing query work. Measure latency and recall together. |
| IVFFlat | Number of lists, number of probes, and whether the index was built after the table contained data | More probes can improve recall at a speed cost. Too many lists for the data available at index-build time can lead to too few results. |
HNSW: expand the candidate search deliberately
The pgvector project README documents a default hnsw.ef_search of 40. Treat this as a configuration default, not a performance target. Increase it in measured increments and observe both query latency and result quality. A broader search may help when the query returns too few useful candidates, but it can cost more time.
Rank #3
For filtered queries, iterative scans can continue beyond the initial approximate scan to find enough matching rows. The README documents hnsw.iterative_scan = strict_order and relaxed_order. Strict ordering keeps exact distance order; relaxed ordering permits small distance-order deviations and can improve recall. Iterative scans are bounded by hnsw.max_scan_tuples and available scan memory, so check those constraints if results remain short.
IVFFlat: validate lists and probes against your data
The project’s starting heuristic for lists is roughly rows divided by 1,000 up to one million rows, and the square root of row count above one million. It suggests beginning with probes around the square root of the list count. These are project rules of thumb, not benchmark findings; test them with your corpus and workload. Increasing probes searches more lists and can improve recall at a speed cost. Also confirm the index was created after the table had data: too few rows relative to the number of lists can result in fewer returned neighbors.
Compare index families using the costs that matter
The project documentation describes HNSW as offering a better speed-recall trade-off, with higher build-time and memory costs. IVFFlat builds faster and uses less memory, with a lower query-performance trade-off. Treat that as directional guidance from the project, not a guarantee for a particular dataset. Compare alternatives using representative query latency, buffer reads, recall, build time, memory footprint, and maintenance needs.
5. Diagnose filters and tenant layout
Approximate-index filtering happens after the vector index scan. That can leave too few results even when the vector index itself is working as designed. In the README’s illustration, a filter matching 10% of rows combined with the default HNSW search breadth of 40 yields about four matching rows on average. This is an explanatory estimate, not a benchmark or a promise for an individual query.
Choose an approach based on filter selectivity
- Highly selective filters: Try an ordinary index on the filter column and exact nearest-neighbor search over the smaller candidate set.
- Approximate search with filters: Try iterative scans so the scan can continue looking for matching rows, subject to the configured scan limits.
- A few distinct filter values: Consider a partial vector index per value when that design fits the schema and query patterns.
- Many values or tenant isolation: Consider partitioning. The project warns that tenants sharing one approximate index can affect one another’s recall and speed; list partitioning or separate tables are named isolation approaches.
Evaluate the filter and tenant pattern alongside the vector index. A faster global vector scan is not a win if some tenants consistently receive too few results or suffer different latency.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.6. Reduce the working set and keep indexes maintained
Consider lower-footprint representations at scale
The pgvector README suggests halfvec for a smaller working set and binary quantization with reranking to help keep indexes in memory at scale. These choices can change precision or search behavior. Compare their results against the current setup and the application’s accuracy requirements before adopting them.
Check HNSW vacuuming and index-build progress
If HNSW vacuuming is slow, the project suggests running REINDEX INDEX CONCURRENTLY before VACUUM. Use the actual index name in the reindex command, and schedule maintenance with the operational impact of the deployment in mind:
REINDEX INDEX CONCURRENTLY your_hnsw_index;
VACUUM your_table;
To see index-build progress, the README shows inspecting pg_stat_progress_create_index, with progress queries for HNSW and IVFFlat. Use those phase details to distinguish a slow build from a query-time problem.
7. Use a controlled troubleshooting loop
- Record the baseline: Capture the representative query, parameters,
EXPLAIN (ANALYZE, BUFFERS)output, latency, rows returned, and result quality against exact search where feasible. - Check query compatibility: Use ascending distance-operator ordering with
LIMIT; verify the plan rather than assuming an index is in use. - Isolate the search mode: Determine whether exact scanning is the cost problem or an approximate index is producing too few or lower-quality neighbors.
- Change one relevant factor: Adjust HNSW breadth or iterative scanning, IVFFlat lists or probes, filter indexing, or representation—one experiment at a time.
- Re-measure the same workload: Compare latency, buffer reads, recall, result count, memory/build implications, and tenant behavior before keeping the change.
pgvector’s README is on the moving master branch and was accessed October 4, 2026. Defaults and feature availability can differ by installed extension release. Check your deployment’s pgvector version and consult documentation matching that release before relying on version-specific settings.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors

