Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For most semantic-search workloads, start with cosine distance (<=>)—unless your embedding model’s documentation or tests on your own data favor another metric. If your vectors are normalized to length 1, cosine and Euclidean distance produce the same rankings, and pgvector recommends inner product (<#>) for best performance on normalized vectors. Whichever metric you choose, match the query operator to the index operator class and compare approximate-search results with an exact-search baseline.

What each pgvector distance operator means

pgvector’s nearest-neighbor operators return distances, so the usual query sorts in ascending order: smaller distance means a nearer result. The operator determines the metric used to rank vectors.

Metric Operator Meaning and use
Euclidean (L2) <-> Straight-line distance between vectors.
Negative inner product <#> Negative dot product, arranged so an ascending index scan can find the largest inner products first. Negate the result to report the inner product.
Cosine distance <=> Cosine-based distance. Report cosine similarity as 1 - (embedding <=> query).
Manhattan (L1) <+> Sum of absolute coordinate differences.
Hamming <~> Distance for binary vectors.
Jaccard <%> Distance for binary vectors.

These operators and their meanings are documented in the pgvector project README. Check the documentation for the exact pgvector version deployed in your database when validating available features and configuration.

How to choose cosine, inner product, or L2

Start with the embedding model’s guidance

OpenAI’s embeddings guide recommends cosine similarity and says, “The choice of distance function typically doesn’t matter much.” That guidance is model-scoped: OpenAI documents that its embeddings are normalized to length 1. For those normalized vectors, cosine similarity and Euclidean distance produce identical rankings, and cosine similarity can be computed with a dot product. Do not assume another model normalizes its output; check the documentation for the exact model you use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
GMKtec AI Mini PC Ryzen Al Max+ 395 (up to 5.1GHz) Mini Gaming Computers
  • EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.

For normalized vectors, pgvector recommends inner product for best performance. Because cosine, Euclidean distance, and inner product rank normalized vectors equivalently, choosing among them may not change which results are nearest; operator and index compatibility still matter.

Choose based on your own relevance judgments

If model documentation does not settle the choice, compare metrics on representative queries and judged relevant results. Assess retrieval quality alongside latency and resource use. A metric that is mathematically convenient is not automatically the best choice for a particular model and search task.

Exact search is the reference for ANN tuning

pgvector performs exact nearest-neighbor search by default, which provides perfect recall. HNSW and IVFFlat are approximate indexes: they can reduce query time, but may return different results from exact search. Treat an approximate index as a speed-versus-recall trade-off, not as a change to the meaning of the distance metric.

Rank #2
AMD Ryzen™ AI Halo - Personal AI Desktop Computer - Developer Platform - Linux OS
  • Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
  • 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
  • AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
  • Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
  • Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.

Build a representative exact-search baseline before tuning. Record result quality against relevance judgments or another task-appropriate measure, the number of results returned, latency, and resource use. Compare approximate results against the same queries and baseline to see what the index parameters actually cost in recall.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Match the query operator to the index operator class

The index operator class must correspond to the operator in the nearest-neighbor query. The main vector mappings are:

Query distance operator Index operator class
<=> cosine distance vector_cosine_ops
<#> negative inner product vector_ip_ops
<-> L2 distance vector_l2_ops
<+> L1 distance vector_l1_ops

For example, if your query orders by cosine distance, create an approximate index using vector_cosine_ops. Using a different operator class does not make the query use the intended metric index. pgvector also documents operator classes for binary-vector search; consult its README for those mappings.

Rank #3
GMKtec EVO-X2 AI Mini PC Ryzen Al Max+ 395 Superchip 128GB LPDDR5X 2TB SSD
  • EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
  • AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
  • AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
  • EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
  • QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.

Choose between HNSW and IVFFlat

Index Documented trade-offs When its characteristics may fit
HNSW Better speed/recall trade-off than IVFFlat in pgvector’s qualitative guidance, but slower builds and higher memory use. It requires no training step and can be created before the table contains data. When its retrieval trade-off is worth the greater build time and memory footprint.
IVFFlat Builds faster and uses less memory than HNSW, with a weaker speed/recall trade-off. It partitions vectors into lists and searches a subset of nearby lists. When quicker builds and lower memory use matter, and its measured recall/latency is sufficient.

These are qualitative trade-offs from pgvector, not a universal measured winner. Compare both against your workload if the choice matters; the better option depends on recall, latency, memory, build time, data loading and updates, and filter behavior.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Tune index parameters against the exact baseline

HNSW

Increase hnsw.ef_search to improve recall at the cost of query speed. The construction parameter ef_construction affects recall as well as index build time and insert speed. Tune using representative queries rather than assuming a setting that works elsewhere will suit your data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

IVFFlat

IVFFlat needs data to build its lists, so pgvector advises creating it after the table contains data. Its documented starting heuristics are:

  • Choose the number of lists as approximately rows/1000 for up to one million rows.
  • Above one million rows, start around the square root of the row count.
  • Start with probes around the square root of the number of lists.

These are starting heuristics, not universal settings. Increasing ivfflat.probes improves recall at a speed cost. Measure the result against exact search on your data.

Check filters and production query shape

With approximate indexes, filtering is applied after the index scan. A query with a selective filter can therefore return fewer rows than requested even when matching rows exist. Validate the actual production query, including its filters, rather than testing only unfiltered nearest neighbors.

Depending on the number and distribution of filter values, pgvector documents iterative scans, partial indexes for a few distinct filter values, or partitioning for many values as approaches to consider. These are alternatives to evaluate against your schema and query patterns, not interchangeable fixes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical metric and index evaluation workflow

  1. Verify vectors. Confirm the embedding model, dimensions, and whether vectors are normalized. Ensure stored and query vectors use compatible model and dimension settings.
  2. Measure exact search. Run exact nearest-neighbor queries on representative evaluation queries. Record relevance quality, result count, latency, and resource use.
  3. Select the metric. Begin with the model’s recommendation; compare alternatives if workload evaluation suggests a benefit.
  4. Build a matching index. Choose HNSW or IVFFlat and use the operator class that matches the query operator.
  5. Tune and compare. Adjust HNSW or IVFFlat settings, then compare approximate results and latency with the exact baseline.
  6. Test realistic filters. Check recall and result counts with production filters and consider the documented filtering approaches where needed.
  7. Monitor over time. Recompare approximate results with exact results as data or query patterns change. The pgvector README shows how to disable index scans in a transaction to obtain an exact comparison.

OpenAI’s guidance on cosine and normalized embeddings is in its official embeddings guide. Both that model-specific guidance and pgvector configuration should be checked against the exact model and extension version in use.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.