Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Hierarchical clustering is an unsupervised machine-learning method that organizes observations into nested groups. It repeatedly merges the most similar clusters (agglomerative clustering) or splits larger groups (divisive clustering), producing a hierarchy that can be inspected at many levels rather than forcing one final number of clusters. A dendrogram shows that hierarchy; its branch heights show the dissimilarity at which merges occur.

What is hierarchical clustering?

Hierarchical clustering builds a tree of nested clusters from unlabeled data. In the common agglomerative, or bottom-up, version, every observation starts as its own cluster. The algorithm repeatedly joins two clusters according to a distance metric and a linkage rule. Divisive methods work in the opposite direction by splitting groups successively.

As scikit-learn explains, “Hierarchical clustering is a general family of clustering algorithms that build nested clusters by merging or splitting them successively.” The result is a hierarchy, not inherently one definitive flat partition. You choose a level of that hierarchy later by cutting the tree or requesting a cluster count.

The result depends on more than the observations themselves. The distance metric, feature scaling and preprocessing, linkage rule, and any connectivity constraints all affect which merges are possible and which groups appear.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Read scikit-learn’s clustering guide.

How agglomerative clustering builds its hierarchy

  1. Start with one cluster per observation.
  2. Compute distances between candidate clusters using the selected linkage rule.
  3. Merge the pair with the smallest resulting inter-cluster distance.
  4. Update the distances involving the new cluster.
  5. Repeat until all observations belong to one hierarchy.

Because each merge contains the clusters below it, the groups are nested: a broad cluster can contain several smaller clusters found at earlier stages.

Linkage methods compared

Linkage defines how the distance between two clusters is calculated. There is no universally best choice; select the rule that matches what “similar” should mean in your application.

Linkage Definition Typical implication Compatibility note
Single Minimum pairwise distance between the two clusters Can follow non-globular structure, but is sensitive to noise and can create uneven, chain-like groups. Available in the documented SciPy and scikit-learn interfaces.
Complete Maximum pairwise distance between the clusters Uses a farthest-pair notion of compactness; assess whether that geometry fits the data. Available in the documented SciPy and scikit-learn interfaces.
Average Mean of all pairwise distances between the clusters Balances the closest- and farthest-pair extremes and is a documented option for non-Euclidean metrics in scikit-learn. Check the selected library’s metric requirements.
Ward Chooses merges that minimize the increase in within-cluster variance Often produces more regular cluster sizes in scikit-learn’s description. Requires Euclidean distance in the documented SciPy and scikit-learn APIs.

These definitions are documented by SciPy’s linkage reference, scikit-learn’s clustering guide, and the AgglomerativeClustering API.

How to read a dendrogram

A dendrogram draws each merge as a U-shaped connector between its child clusters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Leaves: the individual observations, or labels representing them.
  • Branches: the nested relationships created by successive merges.
  • Merge height: the linkage distance at which the two child clusters were joined. A higher merge means greater dissimilarity on the plotted distance scale.
  • Cut: a horizontal threshold that converts the hierarchy into a flat set of clusters.

To choose a grouping, draw an imaginary horizontal line at a selected height. Every disconnected branch below that line is one cluster. Alternatively, request a specified number of clusters from a library. The cut is an analytical decision driven by the task and validation, not an objectively correct answer supplied by the picture alone.

Leaf order can be rearranged for visual readability without changing the hierarchy. Do not interpret neighboring leaves in the horizontal order as a reliable similarity scale; branch structure and merge heights carry the clustering information. SciPy documents these plotting behaviors in its dendrogram reference.

A practical workflow

1. Define meaningful similarity

Specify what it means for two observations to resemble each other. For customers, it might be standardized purchasing behavior; for documents, a vector-space distance; for measurements, a physically meaningful metric. This definition should precede the algorithm choice.

2. Prepare features and distances

Put features on comparable scales when a large-unit variable would otherwise dominate the metric. Handle missing values and extreme outliers using methods appropriate to the domain. You can supply raw observation vectors or a condensed pairwise-distance vector to SciPy’s linkage function.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Metric and linkage compatibility matters. Ward minimizes Euclidean within-cluster variance, so using a non-Euclidean distance with Ward is not valid in the documented implementations. If your similarity is non-Euclidean, compare compatible alternatives such as average, complete, or single linkage.

3. Compare plausible linkages

Run more than one defensible linkage choice, inspect dendrograms, and compare the resulting memberships. A dramatic change between methods is useful evidence that the grouping is sensitive to the geometry rather than an immutable property of the data.

4. Choose a cut for the application

Select a height or cluster count based on the decision the groups must support, interpretability, stability, and any external validation available. A visible gap in merge heights can suggest a candidate cut, but it does not prove that the resulting number of clusters is natural.

5. Validate the groups

Check whether members of each cluster are meaningfully similar under the original domain definition, whether clusters are stable under reasonable preprocessing changes, and whether the partition helps the downstream task. Document the metric, scaling, linkage, cut rule, and date of the analysis so the result is reproducible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Using common software libraries

SciPy

scipy.cluster.hierarchy.linkage builds the hierarchy and dendrogram visualizes it. The linkage function accepts either an observation matrix or a condensed pairwise-distance vector.

from scipy.cluster.hierarchy import linkage, dendrogram, fcluster

Z = linkage(X, method="average", metric="euclidean")
dendrogram(Z)
labels = fcluster(Z, t=4, criterion="maxclust")

Use a method and metric that are mathematically compatible; for example, Ward is defined for Euclidean distances in SciPy’s documented implementation. See the linkage API and dendrogram API.

scikit-learn

AgglomerativeClustering exposes linkage and metric controls, a requested cluster count, and a distance-threshold option. Its documented linkage choices include Ward, single, average, and complete.

from sklearn.cluster import AgglomerativeClustering

model = AgglomerativeClustering(
    n_clusters=4,
    linkage="complete",
    metric="euclidean"
)
labels = model.fit_predict(X)

For Ward, use Euclidean metric as required by the documented API. The current parameter behavior is described in scikit-learn’s AgglomerativeClustering documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

R

R’s stats::hclust performs hierarchical clustering and supports dendrogram output. Its reference documentation is available at R: Hierarchical Clustering.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Computational limits

Hierarchical methods can become expensive because the algorithm considers many possible merges and retains pairwise distance information. SciPy documents O(n²) time for its single, complete, average, weighted, and Ward implementations and O(n²) memory for the described algorithms; some other methods are documented as O(n³) time. These are algorithmic complexity statements, not guaranteed wall-clock benchmarks.

Scikit-learn notes that unconstrained agglomerative clustering considers all possible merges at each step. Connectivity constraints can restrict candidate merges and reduce the search space when the data has a known neighborhood structure. For large datasets, estimate memory requirements before fitting, consider justified connectivity constraints, and evaluate whether a different clustering strategy is more appropriate.

Common interpretation mistakes

  • Treating the dendrogram as a final answer: it represents many possible partitions; the cut still requires a reason.
  • Assuming the tallest branch is automatically meaningful: merge heights depend on the metric, preprocessing, and linkage scale.
  • Reading leaf order as a ranking: rotations of branches can change left-to-right order without changing memberships.
  • Using Ward with an incompatible metric: Ward’s variance objective is Euclidean in the cited implementations.
  • Ignoring scale and outliers: the distance calculation may then reflect units or anomalies rather than the intended concept of similarity.
  • Calling one linkage universally superior: each rule encodes a different geometry and should be judged against the application.

Frequently Asked Questions

Does hierarchical clustering require the number of clusters in advance?

No. It builds the full hierarchy first. You can choose a cluster count afterward or cut at a distance threshold, although that choice still needs application-based validation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use a precomputed distance matrix?

Yes. SciPy’s linkage API accepts a condensed pairwise-distance vector, allowing distances computed with a metric other than the default observation-space calculation. Ensure the chosen linkage supports that metric; Ward requires Euclidean distance in the documented implementation.

Why do two dendrograms from the same data differ?

Differences usually come from preprocessing, distance metric, linkage rule, connectivity constraints, or the library’s input and tie-handling details. The data alone does not uniquely determine a hierarchy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.