Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Changing units can change clustering because distance-based algorithms measure numerical differences, not physical meaning. A feature recorded in metres has numbers 100 times larger than the same feature in centimetres, so it can dominate distances and move observations between clusters. Normalize features before clustering when you need results that are less dependent on measurement scale. Two practical choices are within-feature rank transformation and unit-variance normalization; each solves a different part of the problem.

Why changing units changes a clustering result

Many clustering methods classify observations using distances, similarities or variances. Those quantities depend on the numerical scale of each feature. If one axis is measured in years and another in days, the day-based axis contributes roughly 365 times larger differences even when both describe equally important variation. The same issue appears when converting metres to centimetres, dollars to cents or kilograms to grams.

Rescaling one axis can therefore produce a different apparent arrangement of points and different cluster assignments. That does not prove either arrangement is objectively correct. Vincent Granville’s 9 June 2018 discussion presents the issue as a warning: a visual or algorithmic cluster can be partly an artifact of the coordinate scale.

The manuscript’s random-point illustration reinforces that caution. It uses five points generated with Excel’s RAND() function and argues that repeating the exercise many times can produce apparent clusters in a majority of simulations. This is an illustration of how patterns can appear by chance, not evidence that a particular observed dataset is random or that its clusters have no meaning.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall

Normalize features before clustering

The proposed remedy is to transform each variable before classification so that the algorithm does not give one feature extra influence merely because of its units. The accompanying manuscript describes two approaches.

Rank-transform each feature

Replace every value in a feature with its rank among the observations, processing each feature independently. The smallest value receives the lowest rank and the largest receives the highest. Ties require a stated tie rule, such as assigning average ranks.

Rank #2
Sale
Statistics Laminate Reference Chart: Parameters, Variables, Intervals, Proportions (Quickstudy: Academic )
  • This guide is a perfect overview for the topics covered in introductory statistics courses.

Because ranks preserve order, any monotonic transformation of a feature—such as changing units or applying a monotonic nonlinear function—leaves the ordering unchanged. The resulting representation is therefore insensitive to those transformations. Its trade-off is interpretability: a rank says where an observation stands relative to the sample, not how far it is from another observation in the original units.

Granville characterizes ranks as more robust and less sensitive to noise when distributions are relatively unimodal and do not contain large gaps. That is a qualification, not a universal performance guarantee. Rank transformation can discard meaningful magnitude information, compress large gaps and make distances between adjacent ranks uniform even when the original measurements were not.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3

Normalize each feature to unit variance

Rescale every feature so its variance is one (commonly by subtracting its mean and dividing by its standard deviation). This keeps a continuous magnitude scale while putting features with different spreads on a comparable footing. A linear change of units changes both a feature’s values and its standard deviation by the same factor, so the standardized values are unchanged apart from the usual numerical convention for centering.

Unit-variance normalization is often easier to explain when the original magnitudes remain important. It does not provide the same invariance as ranks to arbitrary monotonic nonlinear transformations, and extreme values can still affect the mean and variance used in the transformation.

Rank transformation versus unit variance

Question Within-feature ranks Unit-variance normalization
What is made comparable? Relative order within each feature Spread, measured by variance, within each feature
Invariance Unaffected by monotonic transformations that preserve order, including unit changes Targets linear scale differences and variance differences
Original magnitude Discarded; values become relative positions Retained as normalized magnitudes
Distribution condition highlighted by the author Favored when distributions are relatively unimodal and lack large gaps No equivalent condition is stated in the cited discussion
Adding new observations Existing ranks may change when ranks are recomputed Means and standard deviations may change when parameters are refit
Main interpretability cost Distances between ranks need not represent original measurement gaps Values are expressed in standard-deviation units rather than domain units

The difficult part: new or streaming data

Rank-based preprocessing depends on the reference dataset. If new training points are added and ranks are recalculated, the transformed values of earlier observations can change. That can alter distances and cluster assignments, so the original clustering structure is not automatically preserved. The manuscript identifies consistent preservation under such additions as a central difficulty.

For a production system, define the reference population and update policy before fitting clusters. You can freeze the transformation used for a deployed model, or deliberately refit both the transformation and clustering when the data distribution changes. Either choice needs monitoring because a frozen reference can become unrepresentative, while refitting can change historical assignments.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Does scaling affect linear-regression coefficients?

Yes, the numerical coefficient can change when a measured quantity is expressed in different units, even when the modeled relationship is unchanged. The manuscript discusses this as a limited scale-invariance property of linear regression: under a linear unit conversion, the attached coefficient is rescaled so predictions represent the same relationship.

Its example changes a coefficient of 3.7 per kilometre to 3.7/1000 per metre. The number changes because the unit changed; it is not evidence that the underlying association or fitted relationship disappeared. Always report the unit with a coefficient.

This statement is limited to linear rescaling. It should not be expanded into a claim that every regression procedure is unaffected by scaling. Logarithmic transformations do not preserve the same coefficient-rescaling property, and predictor scaling can matter for regularized regression, numerical conditioning and coefficient comparisons. The clustering normalization choices above are preprocessing strategies for clustering, not a universal solution for regression.

How to choose a normalization

  1. Identify the algorithm’s sensitivity. Distance-, similarity- or variance-based clustering is scale-sensitive unless the method or implementation explicitly handles scaling.
  2. Define the reference data. Decide which observations determine ranks, means and variances, especially if future data will be assigned to existing clusters.
  3. Use ranks when order is the reliable signal. This is most aligned with the author’s recommendation for relatively unimodal features without large gaps and when invariance to monotonic transformations is important.
  4. Use unit variance when magnitudes still matter. Standardization preserves more information about relative distances while preventing a high-spread feature from dominating solely through its units.
  5. Check sensitivity rather than declaring a “true” clustering. Compare assignments across defensible transformations and inspect whether conclusions depend on one scale choice. Persistence under one transformation is not proof that the resulting structure is objectively correct.

What scale invariance can and cannot promise

  • It can prevent a simple unit conversion from arbitrarily changing a distance-based representation when the chosen transformation has the required invariance.
  • It cannot determine whether a cluster is real, useful or causally meaningful.
  • It cannot remove the effects of outliers, missing data, an unsuitable distance metric or an inappropriate number of clusters.
  • It cannot guarantee stable results when the reference data, ranks or variance estimates are updated.

The chapter on scale-invariant methods in Vincent Granville’s New Statistical Foundations for ML develops these examples, including clustering, regression and spreadsheet computations. The available source does not establish a current retail edition, format or price.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.