Cosine similarity is the dot product of two vectors divided by the product of their L2 norms. Use NumPy for a clear calculation between two dense vectors, scikit-learn for pairwise comparisons and sparse data, and SciPy’s scipy.spatial.distance.cosine only when you want cosine distance: its result is 1 - similarity.
What cosine similarity measures
A vector represents an object as numerical features: a document might be represented by word counts or TF-IDF values, while an image or sentence might be represented by an embedding. Cosine similarity compares the angle between two such vectors, not their lengths. Multiplying a nonzero vector by a positive scalar preserves its direction, so [1, 2, 3] and [2, 4, 6] have a similarity of 1.
For nonzero real-valued vectors, the score ranges from -1 to 1:
- 1: same direction.
- 0: perpendicular vectors.
- -1: opposite directions.
The formula is:
cosine_similarity(a, b) = (a · b) / (||a||₂ × ||b||₂)
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
Here, a · b is the dot product, and the L2 norm is ||a||₂ = sqrt(sum(aᵢ²)). Equivalently, the numerator is sum(aᵢ × bᵢ), divided by each vector’s L2 norm. Scikit-learn describes cosine similarity as the L2-normalized dot product in its metrics documentation.
The score describes the supplied numerical representation, not an object’s meaning in isolation. Similarity between TF-IDF vectors and similarity between learned sentence embeddings reflect different representations and should not be assumed to mean the same thing.
Calculate cosine similarity with NumPy
For two one-dimensional dense vectors, calculate their dot product and norms directly:
import numpy as np
a = np.array([1, 2, 3], dtype=float)
b = np.array([4, 5, 6], dtype=float)
dot_product = np.dot(a, b)
norm_a = np.linalg.norm(a)
norm_b = np.linalg.norm(b)
similarity = dot_product / (norm_a * norm_b)
print(dot_product) # 32
print(norm_a) # 3.741657386...
print(norm_b) # 8.774964387...
print(similarity) # 0.974631846...
The calculation is 32 / (sqrt(14) × sqrt(77)), approximately 0.9746. For one-dimensional arrays, the @ operator is a concise alternative to np.dot:
similarity = (a @ b) / (np.linalg.norm(a) * np.linalg.norm(b))
A reusable function should reject incompatible inputs and handle zero vectors explicitly:
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
import numpy as np
def cosine_similarity(a, b):
a = np.asarray(a, dtype=np.float64)
b = np.asarray(b, dtype=np.float64)
if a.ndim != 1 or b.ndim != 1:
raise ValueError("a and b must be one-dimensional vectors")
if a.shape != b.shape:
raise ValueError("a and b must have the same shape")
if not np.all(np.isfinite(a)) or not np.all(np.isfinite(b)):
raise ValueError("Vectors must contain only finite numeric values")
norm_a = np.linalg.norm(a)
norm_b = np.linalg.norm(b)
if norm_a == 0 or norm_b == 0:
raise ValueError("Cosine similarity is undefined for a zero vector")
score = np.dot(a, b) / (norm_a * norm_b)
return float(np.clip(score, -1.0, 1.0))
Clipping protects against tiny floating-point overshoots such as 1.0000000000000002; it does not fix invalid shapes, NaNs, or incompatible features. The zero-vector exception is one reasonable policy. An application could instead return NaN, exclude empty records, or define a special case such as treating two empty documents as a match. Those are application rules, not values dictated by the formula.
Use SciPy when you need cosine distance
SciPy’s scipy.spatial.distance.cosine returns distance, not similarity. Convert the result by subtracting it from 1:
from scipy.spatial.distance import cosine
a = [1, 2, 3]
b = [4, 5, 6]
distance = cosine(a, b)
similarity = 1 - distance
print(similarity) # approximately 0.974631846
SciPy documents cosine(u, v, w=None) as a distance function for one-dimensional inputs; the optional w supplies weights. For example:
from scipy.spatial.distance import cosine
a = [1, 2, 3]
b = [4, 5, 6]
weights = [1, 2, 1]
similarity = 1 - cosine(a, b, w=weights)
See the SciPy distance reference and the cosine function reference. This API is most direct for a pair of dense vectors or an existing SciPy distance workflow; scikit-learn is generally more convenient for feature matrices and sparse inputs.
Use scikit-learn for vectors and pairwise comparisons
Scikit-learn treats each row as one sample and each column as one feature. Even one vector must therefore be wrapped as a row, giving a two-dimensional input:
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
from sklearn.metrics.pairwise import cosine_similarity
a = [[1, 2, 3]]
b = [[4, 5, 6]]
scores = cosine_similarity(a, b)
print(scores) # [[0.97463185]]
print(scores[0, 0]) # 0.974631846...
The result has shape (number of rows in a, number of rows in b). That is why comparing one row with one row produces a 1 × 1 matrix rather than a scalar. The API accepts sample-by-feature matrices and supports sparse inputs; its parameters and shape requirements are documented in the scikit-learn pairwise implementation reference.
Compare every row with every other row
Pass one matrix to both arguments to get an n × n matrix, where n is the number of rows:
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity
X = np.array([
[1, 0, 0],
[0, 1, 0],
[1, 1, 0],
])
matrix = cosine_similarity(X)
print(matrix)
matrix[i, j] is the score between rows i and j. For nonzero rows, the diagonal should be approximately 1; the matrix should be symmetric, subject to floating-point effects.
Compare two collections
Pass the two collections separately to get a rectangular result. Each collection must have the same number and ordering of features:
documents = np.array([
[1, 0, 1],
[0, 1, 1],
])
queries = np.array([
[1, 1, 0],
])
scores = cosine_similarity(documents, queries)
print(scores.shape) # (2, 1)
Each row of the result corresponds to a document and each column to a query. For large collections, note that a full comparison of n rows with themselves contains n² entries. Compare queries against a collection, batch work, or avoid producing a full matrix if the task needs only a small number of matches.
Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Compare text with TF-IDF
Text must first be represented as vectors. TfidfVectorizer fits a vocabulary and creates a sparse feature matrix; cosine similarity can then compare its rows:
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
documents = [
"Python calculates vector similarity",
"Python calculates cosine similarity",
"Cats sleep on furniture",
]
vectorizer = TfidfVectorizer()
X = vectorizer.fit_transform(documents)
scores = cosine_similarity(X)
print(scores)
To compare a new query with those documents, transform it with the already-fitted vectorizer. Do not fit a new vectorizer for the query: a separately fitted vocabulary may assign different meanings or ordering to feature columns.
query = vectorizer.transform([
"How do I calculate cosine similarity in Python?"
])
scores = cosine_similarity(query, X).ravel()
for document, score in zip(documents, scores):
print(f"{score:.3f} - {document}")
The resulting scores rank these vectors under this TF-IDF representation. They do not establish a universal notion of semantic closeness.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Work with sparse or repeatedly compared vectors
Text feature matrices are often sparse: most documents contain only a small share of the vocabulary. Keep them sparse rather than converting them to dense arrays with .toarray() or .todense(), which can consume substantial memory. Scikit-learn’s metrics documentation describes sparse support for cosine_similarity.
When both inputs are sparse, dense_output=False can request a sparse result:
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchBest Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
scores = cosine_similarity(X, Y, dense_output=False)
For vectors that will be compared repeatedly, normalize each row once and use dot products. Scikit-learn documents that cosine similarity is equivalent to linear_kernel for already L2-normalized data, with the latter avoiding another cosine-normalization step:
from sklearn.preprocessing import normalize
from sklearn.metrics.pairwise import linear_kernel
X_normalized = normalize(X)
scores = linear_kernel(X_normalized, X_normalized)
Both vectors must be normalized for their plain dot product to equal cosine similarity. For a normalized dense matrix, matrix multiplication is another option:
X_normalized = normalize(X)
similarities = X_normalized @ X_normalized.T
For learned embeddings, producing the vectors and comparing them are separate operations. Once a model has produced embeddings, a pair can be compared as rows:
embedding_a = model.encode("A sentence about Python")
embedding_b = model.encode("A sentence about programming")
score = cosine_similarity([embedding_a], [embedding_b])[0, 0]
The score depends on the embedding model, its training and domain, and any preprocessing. Cosine similarity does not itself interpret language.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Choose the right Python API
| Need | Approach | Returns |
|---|---|---|
| Understand or customize the calculation for two dense vectors | NumPy formula | Similarity |
| Use a single-pair SciPy distance workflow or weights | scipy.spatial.distance.cosine, then subtract from 1 for similarity |
Distance |
| Compare rows in dense or sparse feature matrices | sklearn.metrics.pairwise.cosine_similarity |
Similarity matrix |
| Need a pairwise cosine-distance matrix | sklearn.metrics.pairwise.cosine_distances |
Distance matrix |
Scikit-learn defines cosine distance as 1.0 - cosine similarity; see its cosine_distances reference. Keep the distinction clear when naming variables or interpreting output.
Fix common input and interpretation problems
- Different lengths: the vectors must have the same number of features. In scikit-learn, inputs also need the same number of columns.
- Different feature spaces: equal-length arrays are still incompatible if their columns mean different things. Generate text vectors with the same fitted vectorizer, or embeddings with the same model and compatible configuration.
- Zero vectors: the denominator is zero, so the mathematical result is undefined. Detect them and apply an explicit application policy instead of relying on a warning or an assumed score.
- Row versus vector shape: use one-dimensional arrays for a manual NumPy function that expects vectors; use nested rows for scikit-learn.
- Missing or infinite values: reject or handle them before computing. Imputing missing values as zero is not automatically appropriate.
- Negative scores: raw counts and standard TF-IDF vectors are nonnegative, so their scores are generally nonnegative. Embeddings can contain negative coordinates, making negative cosine scores possible.
- Thresholds: a score such as 0.8 is not a universal cutoff for “similar.” Choose thresholds against labeled examples or an evaluation set for the model and task.
Install the libraries and check your environment
Install the packages used in the examples with:
python -m pip install numpy scipy scikit-learn
For an isolated environment, create and activate a virtual environment first:
python -m venv .venv
On macOS or Linux, activate it with source .venv/bin/activate. In Windows PowerShell, use .venvScriptsActivate.ps1. The cited scikit-learn stable documentation and SciPy reference pages may identify specific documentation versions; those labels do not establish what is installed locally. Check the local package versions with:
Quick Recap
python -c "import numpy, scipy, sklearn; print(numpy.__version__, scipy.__version__, sklearn.__version__)"
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

