Choose a nonparametric test by how your observations were collected—not simply because your data fail a normality check. The key distinction is whether groups are independent, paired, or arranged in blocks. Common rank tests can work with ordinal data and avoid a normality model for raw observations, but they still rely on assumptions such as independence and, for some tests, symmetry.
When do we require nonparametric or distribution-free methods?
Nonparametric methods are useful when measurements are ordinal or naturally ranked, when a parametric model’s assumptions are not defensible, or when the question concerns properties such as randomness, independence, symmetry, or goodness of fit. They are not automatically the right choice whenever a normality test is significant: first consider the study design and the quantity you want to learn about.
The terms nonparametric and distribution-free are related, but not perfectly interchangeable. NIST explains that distribution-free procedures have test statistics that do not depend on the form of the underlying distribution, while nonparametric procedures are not concerned with distribution parameters. In practice, many introductory rank tests avoid assuming normally distributed raw observations, but they are not assumption-free. Meaningful ranks, independent observations, or symmetric paired differences may still be required. NIST’s discussion of nonparametric methods outlines these distinctions.
When the assumptions of a parametric method are reasonable, it may be more efficient. Choose a procedure that fits the design and research question, rather than treating either method family as universally safer.
Recommended Free Tools
#1 Best Overall
How can you choose a test from the study design?
Start by asking whether the observations come from separate, unrelated groups or from the same units measured more than once. If measurements are matched or grouped into blocks, preserve that structure in the test.
| Design | Common choice | What the test does |
|---|---|---|
| Two independent groups | Mann–Whitney U, also called Wilcoxon rank-sum | Ranks the pooled observations and compares rank behavior between groups. It is not a paired test; ties receive average ranks in NIST’s handbook description. |
| More than two independent groups | Kruskal–Wallis | Ranks the pooled observations and compares group rank sums. A significant omnibus result does not identify which groups differ. |
| Two paired conditions or matched observations | Wilcoxon signed-rank | Calculates within-pair differences, ranks their absolute values, then restores their signs. It assumes the differences are mutually independent and symmetric. |
| Paired observations when difference magnitudes should not be used, or symmetry is doubtful | Sign test | Uses only the direction of nonzero paired differences. Because it ignores magnitude, it makes fewer assumptions than signed-rank. |
| Several treatments measured in blocks or on the same experimental units | Friedman | Ranks treatments within each block. Blocks should be mutually independent, and measurements must be meaningfully rankable. |
NIST describes the Mann–Whitney test, Kruskal–Wallis test, signed-rank and sign tests, and the Friedman test in more detail.
Rank #2
- This guide is a perfect overview for the topics covered in introductory statistics courses.
What does a rank-test result mean?
Mann–Whitney U and Wilcoxon rank-sum
These names refer to closely related ways of comparing two independent groups using pooled ranks. Interpret the result according to the procedure and null hypothesis used by your software. Although the test is sometimes introduced as a comparison of central tendency, it can respond to distributional differences more broadly. Do not automatically report a significant result as a difference in medians: that interpretation requires suitable conditions on the distributions, such as similar shapes.
Kruskal–Wallis
A Kruskal–Wallis rejection is evidence against the hypothesis that all groups have the same distribution or rank behavior under the test setup. It does not tell you which pairs differ. Use an appropriate follow-up comparison and account for the multiplicity created by making several comparisons.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
The usual chi-square approximation for the H statistic is not equally suitable for every sample size. NIST’s handbook describes H as approximately chi-square with k−1 degrees of freedom when group sizes are not too small, giving ni > 4 as a rule of thumb. NIST Dataplot states that each group should have at least 5 observations for its approximation. These are guidance from the respective NIST references, not universal guarantees. For very small samples, use an exact or otherwise suitable method supported by the software you use. See NIST’s Kruskal–Wallis handbook page and its Dataplot reference.
Wilcoxon signed-rank and sign tests
For signed-rank, the unit of analysis is the within-pair difference, not the two sets of measurements treated as independent groups. The test uses both the sign and magnitude of nonzero differences and assumes their distribution is symmetric. The sign test uses only whether each nonzero difference is positive or negative, so it is an option when using magnitudes is not appropriate or symmetry is doubtful. NIST describes signed-rank as relaxing the paired t-test’s normality assumption to symmetry, while the sign test requires fewer assumptions because it ignores magnitude. NIST Dataplot’s signed-rank reference explains both tests.
Rank #4
Friedman
Friedman is for blocked or repeated-condition designs: rank the treatments within each block, rather than pooling every observation across all treatments. A significant omnibus result indicates evidence of a difference among treatments, but does not identify which ones differ; follow it with suitable comparisons that account for multiplicity. The test relies on independent blocks and measurements that can be ranked meaningfully. See NIST Dataplot’s Friedman reference.
Quick Recap
Best Value
Common mistakes to avoid
- Using an independent-groups test on paired data. Pairing carries information; use a test based on within-pair differences or the relevant blocked design.
- Calling a rank-test result a median difference by default. If distributions differ in spread or shape, a rank difference need not mean a median shift.
- Assuming nonparametric means assumption-free. Check independence, rankability, and symmetry where relevant.
- Treating an omnibus result as a pairwise answer. Kruskal–Wallis and Friedman indicate an overall difference; follow-up comparisons are needed to locate it.
- Applying a large-sample approximation without checking sample size. Consider an exact or other appropriate method when groups are very small.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

