Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Euclidean and Manhattan distance measure coordinate differences; cosine distance measures the angle between vectors. Minkowski distance is the broader family that includes Euclidean and Manhattan as special cases. The right choice depends on what “similar” means for your data—and on whether feature scales or vector magnitude should affect the result.

How the four distance measures differ

For two vectors x and y with n coordinates, each measure turns their differences into a single value. A smaller value generally means the vectors are closer under that measure, but the meaning of “closer” changes with the formula.

Measure Definition What it emphasizes
Euclidean (L2) d(x,y) = √Σi(xi − yi)² Straight-line separation. Squaring coordinate differences makes large deviations count disproportionately.
Manhattan (L1) d(x,y) = Σi|xi − yi| The total absolute difference across coordinates; also called city-block distance. Scikit-learn identifies its implementation as L1 distance (scikit-learn Manhattan distances).
Minkowski (Lp) dp(x,y) = (Σi|xi − yi|p)1/p, usually for p ≥ 1 A family controlled by p: p=1 gives Manhattan; p=2 gives Euclidean. Larger coordinate differences have more influence as p increases. Scikit-learn lists Minkowski among its supported pairwise metrics (scikit-learn pairwise distances).
Cosine distance 1 − (x · y)/(||x|| ||y||) Angular dissimilarity: it focuses on vector orientation rather than the absolute magnitudes of coordinates.

How to choose a distance measure

Use Euclidean for geometric closeness

Choose Euclidean when straight-line separation in numeric features reflects the similarity you care about. It can be sensitive to a large difference in one coordinate because that difference is squared before the coordinates are summed.

Use Manhattan for additive coordinate differences

Manhattan distance adds the absolute difference in each coordinate. It is appropriate when a total of per-feature deviations better represents “far apart” than straight-line geometry does.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Minkowski when you need a tunable family

Minkowski lets you set p to control how strongly large coordinate differences affect the result. It is not a wholly separate peer to Euclidean and Manhattan: those are its p=2 and p=1 cases.

Use cosine when direction or pattern matters more than size

Cosine distance is useful when the orientation or relative pattern of vector values carries more meaning than their overall magnitude. For unit-normalized samples, scikit-learn states that cosine distance is half the squared Euclidean distance (scikit-learn cosine similarity and distance). The direction of a zero vector is undefined, so check how your implementation handles zero vectors before using cosine distance.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Scale features before comparing distances

Distance calculations can be dominated by features with larger units or ranges. For heterogeneous numeric data, standardize or otherwise scale features when those differences would otherwise give one feature undue influence. Scaling changes the geometry, so base it on the meaning of the variables and the task rather than applying it mechanically.

For sparse text or embedding vectors, cosine is often considered because direction can carry the useful signal. That is a starting point, not a guarantee of better results: assess the choice against your data and objective.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Distance, similarity, and metric requirements

A distance-like score is not automatically a mathematical metric. Scikit-learn’s user guide says a true metric must be nonnegative, equal zero only for identical objects, symmetric, and satisfy the triangle inequality (scikit-learn metrics, affinities, and kernels). Similarity kernels are a distinct concept; an algorithm that requires a metric may not accept every similarity or distance score interchangeably.

Cosine distance is useful for many vector comparisons, but do not assume it meets every requirement imposed by a metric-based algorithm. Check the algorithm’s documentation and the chosen implementation’s behavior.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Calculating pairwise distances with scikit-learn

The pairwise_distances API computes distances between rows of feature arrays. With Y=None, it returns pairwise distances among rows of X; with metric="precomputed", it accepts a precomputed distance matrix. Its listed options include cosine, Euclidean, Manhattan (L1), and Minkowski (scikit-learn pairwise distances API).

The Euclidean API uses the identity √(x·x − 2x·y + y·y), which can be useful with sparse arrays and precomputed norms. Scikit-learn cautions that this calculation can suffer catastrophic cancellation and that floating-point results may not produce an exactly symmetric distance matrix (scikit-learn Euclidean distances API). Consult the documentation for the installed version when implementing distances, since API details can evolve.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical selection checklist

  • Choose Euclidean if straight-line separation in appropriately scaled numeric features matches your definition of closeness.
  • Choose Manhattan if the sum of absolute coordinate deviations is more meaningful.
  • Choose Minkowski if you need to tune how large coordinate differences contribute, and select p deliberately.
  • Choose cosine if vector direction or relative pattern matters more than magnitude; confirm zero-vector handling.
  • Scale heterogeneous features when units or ranges would otherwise dominate.
  • Confirm that the target algorithm supports the selected metric, then validate the choice against your task rather than assuming one measure is universally best.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.