Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A distance function measures how dissimilar two represented observations are: a smaller value means the selected rule considers them more alike. A metric is a distance function that also satisfies four mathematical properties, so not every function used to compare machine-learning data is strictly a metric. Which one to use depends on what the features represent, whether magnitude matters, and what the algorithm accepts.

What is a distance metric?

A distance function assigns a value to a pair of observations, often represented as feature vectors. Scikit-learn describes the comparison this way: “Distance metrics are functions d(a, b) such that d(a, b) < d(a, c) if objects a and b are considered “more similar” than objects a and c.” The value is meaningful only in relation to the chosen representation and comparison rule; it has no context-free interpretation.

In code and informal discussion, “distance” is sometimes used for any dissimilarity score. A strict mathematical metric has additional requirements. That distinction matters when an algorithm relies on metric properties, or when interpreting what a reported value guarantees.

What makes a distance function a true metric?

For observations x and y, a metric d must satisfy all four conditions:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
  • Non-negativity: d(x, y) is never less than zero.
  • Identity of indiscernibles: d(x, y) equals zero if and only if x and y are identical.
  • Symmetry: d(x, y) = d(y, x).
  • Triangle inequality: d(x, z) ≤ d(x, y) + d(y, z).

A function that violates one or more conditions may still be useful as a dissimilarity for a particular method, but it should not be called a strict metric without qualification. For example, the common cosine distance defined as one minus cosine similarity does not satisfy all metric axioms in general.

How do common distance metrics differ?

For numeric feature vectors, distance is a geometric choice: it defines what counts as close. Minkowski distance is a family that includes Manhattan and Euclidean distance as special cases. Other choices are more appropriate when direction, feature relationships, or a particular data type matters.

Choice How it compares observations Useful when Important qualification
Minkowski Uses the p-norm across vector coordinates. A general numeric-vector distance is appropriate. p = 1 gives Manhattan distance; p = 2 gives Euclidean distance. Scaling of coordinates affects the result.
Manhattan (L1, City Block) Sums the absolute coordinate differences. Coordinate-by-coordinate deviations are the intended comparison. Feature units and scales still affect the total.
Euclidean (L2) Measures straight-line separation between vectors. Ordinary geometric separation in numeric feature space is meaningful. Large-scale coordinates can dominate the result.
Cosine similarity / cosine distance Compares the angle between vectors after L2 normalization; cosine distance is an available pairwise-distance option. Direction or relative pattern matters more than raw magnitude, such as with TF-IDF document vectors. Cosine similarity is not a distance: higher means more alike. One minus cosine similarity is not generally a strict metric.
Mahalanobis Applies a positive semidefinite matrix, equivalently a linear transformation followed by Euclidean distance. Feature relationships or covariance should alter the effective geometry. The transformation and its estimation determine which differences count.
Hamming or Jaccard Compare binary or set-like data according to their respective mismatch or overlap rules. Features are indicators or sets rather than ordinary continuous measurements. Confirm that the chosen definition matches the representation and library support.
Haversine Compares geographic latitude/longitude points using spherical geometry. Coordinates represent geographic locations. The referenced scikit-learn DistanceMetric API specifies radians for inputs and outputs; verify coordinate order and the library version in use.

L1, L2, and Minkowski

Minkowski distance is a parameterized family. Setting p = 1 gives Manhattan (L1) distance, the sum of absolute coordinate differences. Setting p = 2 gives Euclidean (L2) distance, the straight-line distance. The choice is not merely a formula preference: it encodes how coordinate-wise differences combine into an overall separation.

These distances are sensitive to feature scales. If one numeric feature ranges over thousands while another ranges between zero and one, the larger-scale feature can dominate a raw distance. Consider units, normalization, or other preprocessing as part of the modeling decision. There is no single scaling recipe established for every task.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cosine for direction rather than magnitude

Cosine similarity is the dot product of two vectors after L2 normalization. It emphasizes direction: vectors pointing similarly can score as similar even if their original magnitudes differ. That is useful for document vectors when the pattern of term weights is more relevant than document length. Scikit-learn notes that for normalized TF-IDF vectors, cosine similarity is equivalent to the linear kernel; its pairwise API also provides cosine distance. These are related but differently oriented quantities: similarity generally increases with likeness, while distance decreases.

For a focused treatment of vector-space similarity and TF-IDF, see Introduction to Information Retrieval.

Mahalanobis distance and learned geometry

Mahalanobis distance changes the geometry according to a positive semidefinite matrix, or equivalently compares points after a linear transformation using Euclidean distance. This can account for relationships among features instead of treating every coordinate as independent and equally scaled.

Metric learning fits such a transformation using supervision. Depending on the method, that supervision may be labels or relationships such as similar and dissimilar pairs or triplets. The learned geometry can bring related examples closer and push unrelated examples farther apart. If the transformation collapses distinct points to zero distance, it yields a pseudometric rather than a strict metric.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How do you choose a distance metric?

Choose from the meaning of the comparison, not from a list of popular names. Use this checklist before selecting a metric for a model or analysis:

  1. Identify the representation. Decide whether the data are continuous numeric vectors, sparse text vectors, binary indicators, sets, or geographic coordinates. A distance that suits dense numeric features may not suit the others.
  2. Decide whether magnitude should matter. If vector length is meaningful, a magnitude-sensitive distance may fit. If the pattern or direction matters more than length, cosine similarity may be a better starting point.
  3. Check scale and feature relationships. Review units and ranges before applying L1 or L2 to raw features. If feature correlations should affect comparisons, consider a Mahalanobis-style transformation.
  4. Check the algorithm’s mathematical requirements. Determine whether the downstream method requires a true metric or accepts a more general dissimilarity. Do not assume every function called a distance satisfies the metric axioms.
  5. Verify software and input support. Check the documentation for the library version you will run, the accepted metric names, and supported input formats. Scikit-learn’s pairwise options include Euclidean, cosine, Manhattan/City Block, Minkowski, Mahalanobis, Hamming, Jaccard, and others; in the referenced API, not every SciPy-provided metric supports sparse matrices.
  6. Use supervision only when it is available and appropriate. If you have labels or meaningful similar/dissimilar relationships, metric learning can fit a task-specific transformation. Without such supervision, choose and validate a distance based on the feature semantics and modeling objective.

How are distances used in machine-learning code?

Scikit-learn’s pairwise utilities calculate distances between rows of sample matrices and allow an explicit metric argument. The available choices and input constraints depend on the particular API and release, so consult documentation for the version deployed. A distance matrix is not automatically a similarity matrix: distances typically get smaller as observations become more alike, while similarities typically get larger. Any conversion between them should name the transformation and account for its mathematical properties.

For geographic data, the referenced scikit-learn DistanceMetric documentation specifies Haversine inputs and outputs in radians. Check the installed release’s API documentation and confirm coordinate ordering before passing latitude and longitude values.

Is there one best distance metric?

No. The same pair of observations can look close under one representation or rule and far apart under another. A useful choice is the one whose geometry reflects the feature meanings, handles scales and relationships appropriately, matches the algorithm’s requirements, and is supported for the data format in the software version being used.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.