Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

K-means is an unsupervised learning algorithm that partitions numerical observations into a number of groups you choose. It assigns each observation to the nearest cluster center, recalculates centers as means, and repeats. It does not discover a uniquely true number of groups: you choose k, then assess whether the resulting partition is stable, well separated, and useful for your problem.

What k-means clustering does

K-means represents each cluster by its centroid—the mean of the observations assigned to it. For a chosen k, the algorithm alternates between two operations:

  1. Assignment: assign each observation to its nearest centroid.
  2. Update: recompute each centroid as the mean of the observations currently assigned to it.

It repeats these steps until a stopping condition is reached. The objective, called inertia in scikit-learn, is the sum of squared distances from each sample to its closest cluster center. The scikit-learn KMeans documentation describes the algorithm and its objective.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Because assignment is based on distance to centroids, the features you include and their units determine what “near” means. A feature measured in thousands can dominate another measured in fractions if the numeric scales differ. Choose features that make sense for the question, inspect their ranges, and consider scaling where appropriate; there is no single scaler that is right for every dataset.

How to choose the number of clusters

k is a required input, not a value K-means automatically discovers. Choose a plausible range based on what decisions or descriptions the clusters need to support, then compare candidate results. No single score proves that a particular value is the objectively correct one.

Compare inertia, but do not optimize it alone

Fit candidate values of k and compare inertia. Adding clusters will generally reduce this objective, so a lower value by itself is not evidence that the added complexity is useful. Look for a trade-off that suits the use case rather than assuming the lowest inertia is best.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Use silhouette as a separation diagnostic

Silhouette analysis offers another way to inspect candidate partitions. Its score ranges from -1 to 1: values near +1 suggest a sample is well separated from neighboring clusters, values near 0 suggest it lies near a boundary, and negative values may indicate that it has been assigned to a less suitable cluster. The scikit-learn silhouette example shows how to examine scores for candidate clusterings. Treat the score as a diagnostic, not proof of a meaningful taxonomy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect the clusters themselves

Check cluster sizes, representative observations, and whether the groups make sense for the task. A clustering can score well geometrically yet be too imbalanced, hard to explain, or irrelevant to a decision. Compare these observations across candidate values of k before settling on a result.

A practical K-means workflow

  1. Select and inspect features. Use meaningful numerical variables, check their ranges and units, and decide whether scaling is needed so that distance reflects the intended notion of similarity.
  2. Set candidate values of k. Define a reasonable range from the problem, not from a claim that the data contain an inherently fixed number of natural groups.
  3. Fit reproducibly. Control the random state and use enough initializations to check whether results depend on the starting centers. Record preprocessing, library version, parameters, and seed.
  4. Compare the outputs. Review inertia and silhouette alongside cluster sizes and representative observations. Prefer a partition that is both interpretable and useful, not merely one with a favorable metric.
  5. Check whether the geometry fits. If the groups are elongated, density-based, strongly unequal in variance, or distorted by outliers, reconsider whether K-means is an appropriate model rather than trying only more values of k.

Initialization, defaults, and repeatability

Different starting centroids can lead K-means to different local minima and therefore different partitions. Multiple initializations help reveal this sensitivity; a fixed random state makes a run repeatable, but does not establish that its result is best.

In the scikit-learn 1.9.1 KMeans API, documented defaults include n_clusters=8, init='k-means++', and n_init='auto'. With n_init='auto', the API specifies one run for k-means++ or explicit initial centers, and ten for random initialization or a callable. These are library defaults, not general recommendations; choose settings deliberately and check sensitivity for your data.

When K-means may not fit

K-means is most natural when centroid-based, roughly circular groups are a reasonable description of the data. Its distance-based objective can be a poor match for outliers, density-based clusters, anisotropic (elongated or differently oriented) groups, or clusters with unequal variance. The scikit-learn demonstration of K-means assumptions illustrates cases where these shapes challenge the method. If the observed structure conflicts with those assumptions, compare another clustering family instead of forcing K-means to explain it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Scaling to larger datasets

MiniBatchKMeans updates centers using small batches rather than processing the full dataset for every update. Scikit-learn presents more than 10,000 samples as an example scale at which it may be much faster than conventional K-means, not as a guaranteed crossover point. See the scikit-learn MiniBatchKMeans documentation and benchmark both approaches on your data and hardware; speed and result quality depend on the workload.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.