What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For two nonzero vectors with the same features in the same order, cosine similarity is their dot product divided by the product of their Euclidean lengths. Use a short NumPy function for one pair of dense vectors, or scikit-learn’s pairwise function for collections of rows and sparse text data.
What cosine similarity measures
Cosine similarity compares the direction of two vectors rather than their raw magnitude:
similarity(a, b) = dot(a, b) / (||a||₂ × ||b||₂)
Scikit-learn defines it as the L2-normalized dot product. For ordinary real-valued vectors, the result ranges from -1 to 1. With nonnegative features, such as counts or TF-IDF weights, it ranges from 0 to 1. Multiplying a nonzero vector by a positive constant does not change its cosine similarity, so cosine can discard magnitude information that may matter for your task. Scikit-learn’s metrics documentation explains the definition and use.
#1 Best Overall
Implement cosine similarity for one pair of vectors
This NumPy helper checks that both inputs are one-dimensional, have matching shapes, and are nonzero before calculating the normalized dot product:
import numpy as np
def cosine_similarity(a, b):
a = np.asarray(a, dtype=float)
b = np.asarray(b, dtype=float)
if a.ndim != 1 or b.ndim != 1:
raise ValueError("a and b must be one-dimensional vectors")
if a.shape != b.shape:
raise ValueError("a and b must have the same shape")
norm_a = np.linalg.norm(a)
norm_b = np.linalg.norm(b)
if norm_a == 0 or norm_b == 0:
raise ValueError("cosine similarity is undefined for a zero vector")
return float(np.dot(a, b) / (norm_a * norm_b))
For example, [1, 0] and [0, 1] have a dot product of zero, so their cosine similarity is 0. The explicit checks make invalid or incompatible inputs fail clearly instead of producing a misleading score.
Rank #2
Compare rows with scikit-learn
For pairwise comparisons between sets of vectors, use scikit-learn’s cosine_similarity. It returns a matrix whose entry at row i, column j is the similarity between X[i] and Y[j].
from sklearn.metrics.pairwise import cosine_similarity
scores = cosine_similarity(X, Y)
The function accepts SciPy sparse matrices, making it suitable for sparse feature representations such as text vectors. See the pairwise metrics API documentation for the accepted inputs and return shape.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →For text features
Cosine similarity operates on vectors, not raw strings. First convert documents into vectors in the same feature space—for example, with TF-IDF. Scikit-learn notes that when TF-IDF rows are L2-normalized, their dot product is already cosine similarity. Its preprocessing guide describes the normalized-vector shortcut.
For repeated queries against a fixed collection
If the collection’s rows are L2-normalized, compute their dot products with each query to obtain cosine scores. For a matrix of normalized rows, matrix multiplication can calculate many scores at once. This shortcut is valid only when both sides are normalized in the same way; keep that invariant explicit so normalized and unnormalized vectors are not mixed.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Handle edge cases and interpret scores correctly
Zero vectors
A zero vector has a zero norm, making the denominator zero; ordinary cosine similarity is undefined. Reject it, as the helper above does, or define a documented application-specific convention. Do not add an arbitrary epsilon and present the resulting value as the standard formula. Scikit-learn’s normalization implementation handles zero norms internally, but its exact output policy should be checked against the installed release’s documentation if it matters to your application; the main-branch implementation is mutable.
Matching dimensions and feature order
Equal vector lengths are not enough: each coordinate must mean the same feature on both sides. Ensure the vectors share a feature space and ordering. A shape check can catch different lengths, but it cannot detect a semantic mismatch between coordinates.
Best Value
Negative coordinates and magnitude
Negative coordinates can produce negative cosine scores when vectors point in opposing directions. The commonly expected 0-to-1 range applies to nonnegative features, not all real vectors. Also, cosine ignores positive rescaling: if the size of a vector carries useful information, compare with a metric or method that preserves magnitude rather than relying on cosine alone.
Quick Recap
Choose the implementation for your workload
- One pair of small, dense vectors: use the NumPy helper to keep the calculation explicit and easy to inspect.
- Many rows or sparse text features: use scikit-learn’s
cosine_similarity(X, Y)to get pairwise scores. - Rows already L2-normalized: use a dot product or matrix multiplication, provided the normalization is guaranteed for both inputs.
- Embedding vectors: the same formula applies, but whether cosine is appropriate depends on the embedding model and downstream task. A similarity score is not automatically a calibrated probability or universal judgment of meaning.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

