The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
You do not need model-generated embeddings to use pgvector. It can rank any supported numeric vector, including one you build from structured data. A hand-built feature vector is a strong fit when you can name and measure the attributes that should make records similar; embeddings are usually more natural for unstructured text or images whose meaningful features are hard to specify. If the task boils down to a couple of numeric conditions, ordinary SQL may be simpler than either.
Do you need embeddings to use pgvector?
No. pgvector is a PostgreSQL extension for storing vectors and searching by distance. It does not decide how a vector is created or what its dimensions mean. A vector may come from a machine-learning model, or your application can calculate it from database fields.
That distinction matters: pgvector ranks the representation you provide, while your feature design determines what “similar” means. The extension supports several vector types, distance operators, exact search, and approximate indexes. Your choice of dimensions, scaling, weights, and missing-value treatment is part of the application’s relevance logic.
When does a feature vector make more sense than semantic search?
Consider hand-built features when records are structured and the attributes that define similarity are known and measurable. For instance, a product may need to find other products with similar dimensions, capacity, price band, and power use. Those named dimensions make the comparison inspectable and give you direct control over its behavior.
#1 Best Overall
Embeddings are often a better starting point when the input is prose or imagery and you cannot readily enumerate the signals that make two items alike. A model provides a learned representation, but its dimensions are less directly interpretable than explicit fields. Neither approach is inherently more accurate: the right choice depends on whether the representation matches the task, and should be evaluated against application-specific relevance needs.
How to build a useful feature vector
Choose dimensions that answer the product question
Start by stating what “similar” should mean to users, then select measurable columns that express it. A baseball-pitcher example in a 2026 Agave Information Solutions article uses pitch-type shares, pitch-location means and spreads, velocity averages and ranges where available, and changes in pitch mix by count. It combines and normalizes these aggregates into a 32-dimensional vector. That is an illustration, not a universal recipe or a validated feature set for other domains.
Put different scales on a comparable footing
If raw values with very different ranges go into a distance calculation together, a large-scale dimension can dominate. Standardization such as z-scores, or mapping values to a fixed min–max range, can reduce that effect. Choose the transformation for your data distribution and desired behavior, and validate it rather than assuming a particular normalization is automatically right.
Set weights deliberately
Scaling dimensions lets you make some attributes count more than others. That is useful control, but the weights encode a product or domain judgment; explicit weights are not necessarily correct weights. Evaluate whether the resulting neighbors match the intended notion of similarity.
Represent missingness honestly
A missing reading is not automatically a zero. The pitcher example’s author notes that velocity measurements were often missing in that dataset and suggests imputing a population mean or dropping a dimension and renormalizing. Those are possible approaches, not independently validated rules. Pick a strategy that preserves the meaning of missing data in your application.
Choose between feature vectors, embeddings, both, and SQL
| Approach | Consider it when | Main trade-off |
|---|---|---|
| Hand-built feature vector | Records are structured and the useful similarity dimensions are known and measurable. | Feature choice, scaling, weighting, and missing-data behavior define relevance and require task-specific evaluation. |
| Model embedding | Inputs are unstructured, such as prose or images, and meaningful dimensions are difficult to specify by hand. | The representation is learned rather than a list of named, directly interpretable features. |
| Both | Structured attributes and unstructured content each contribute distinct similarity signals. | You must decide how to combine signals; there is no generally established fusion method or guaranteed gain. |
| Ordinary SQL | One or two numeric criteria or straightforward predicates capture the task. | A vector representation and index may add complexity without helping express the query. |
Compare candidates on the dimensions that matter for the application: input structure, whether useful features are known, interpretability, missing-data behavior, and measured relevance. For indexed search, also compare recall, latency, index-build time, and memory under representative data and filters.
Rank #3
What pgvector search and indexes change
The pgvector project says exact nearest-neighbor search is the default and provides perfect recall. HNSW and IVFFlat are approximate indexes: they can improve speed while returning results that differ from exact search. The project describes HNSW as offering a stronger query-speed/recall trade-off than IVFFlat, at the cost of slower index builds and greater memory use. IVFFlat partitions vectors into lists and searches selected lists; it builds faster and uses less memory, with lower query performance in that documented trade-off. These are project descriptions, not workload guarantees.
Free tools Windows power users keep installed
One-click scans. No signup required.
The distance operator and index operator class need to match the intended metric. pgvector documents L2 distance with <->, negative inner product with <#>, cosine distance with <=>, L1 with <+>, and Hamming/Jaccard for binary vectors with <~> and <%>. The negative inner-product operator returns a negative value so it can support ascending index scans.
For IVFFlat, the project recommends creating the index after the table contains data because the index has a training step. Its initial tuning heuristics are rows divided by 1,000 for list count up to one million rows, and the square root of row count above one million. A suggested initial probe count is the square root of the number of lists; more probes generally improve recall at a speed cost. These are starting points, not benchmark findings. Compare indexed results and latency with exact search on your workload.
Rank #4
Account for filters and hybrid retrieval
With approximate indexes, filtering happens after the index scan. The pgvector README illustrates the effect: with a filter matching 10% of rows and HNSW’s default ef_search of 40, an average of four qualifying rows is expected from that scan. If filters or tenant boundaries matter, the project documents iterative scans, indexes on filter columns, partial indexes for a few distinct values, and partitioning for many values as possible approaches. Which is appropriate depends on selectivity and the desired result count, so measure it.
A product that needs both structured similarity and semantic relevance can retain separate signals rather than forcing all information into one representation. The pgvector documentation shows full-text search combined with vector-related search and mentions Reciprocal Rank Fusion or a cross-encoder for combining results. A feature-vector article separately proposes pairing structured features with embeddings for prose. These are design options; neither source establishes a universally best fusion strategy.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteTest the representation before tuning the index
Index tuning cannot repair a vector that encodes the wrong idea of similarity. First inspect whether nearest neighbors make sense to users, including edge cases involving missing values, different scales, and important filters. Then compare exact search as a recall baseline with the approximate index and settings you plan to deploy. Measure relevance and latency on representative queries, and track build time and memory where those affect operations.
Best Value
The baseball example shows the basic SQL pattern; its 32 dimensions and limit of 10 are example-specific, not requirements:
CREATE EXTENSION IF NOT EXISTS vector;
ALTER TABLE pitcher_profiles
ADD COLUMN feature_vec vector(32);
CREATE INDEX ON pitcher_profiles
USING hnsw (feature_vec vector_cosine_ops);
SELECT id, name
FROM pitcher_profiles
WHERE id <> @target_id
ORDER BY feature_vec <=> @target_vec
LIMIT 10;
The vector length, operator class, and query should reflect your own representation and distance choice. Excluding the target row prevents a self-match in this example.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

