Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For most day-to-day analysis, five test families cover a useful first set of questions: t-tests for means, chi-square tests for categorical associations, ANOVA for comparing means across groups, correlation tests for numeric association, and the Mann–Whitney U test for rank-based comparisons of two independent groups. The right choice depends on the outcome, study design, and quantity you want to estimate—not just whether a normality test is significant.

This is a practical shortlist, not a universal canon. A regression model, exact test, or method for repeated or clustered observations may be better than any of these five.

Start with the question and the data

A hypothesis test evaluates how compatible observed data are with a specified null model. The null hypothesis, H0, commonly says there is no difference or association; the alternative, HA, describes the effect or relationship being considered. A test statistic summarizes the data for that question, and its sampling distribution under the null model is used to calculate a p-value.

A p-value is the probability, assuming the null model and test assumptions hold, of obtaining a result at least as extreme as the one observed. It is not the probability that the null hypothesis is true, the probability that the result happened by chance, or a measure of how large or useful an effect is. A small p-value can accompany a trivial effect in a very large sample; a non-significant result in a small sample can leave substantial uncertainty.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Statistics Laminate Reference Chart: Parameters, Variables, Intervals, Proportions (Quickstudy: Academic )
  • This guide is a perfect overview for the topics covered in introductory statistics courses.

Choose a significance threshold, α, before interpreting results when possible. A Type I error is rejecting a true null hypothesis; a Type II error is failing to reject a false one. Statistical power is the probability of detecting a specified effect under stated assumptions. Report an effect estimate and confidence interval alongside the p-value: the interval communicates plausible values under the analysis, while the effect size helps describe magnitude.

Use “reject the null hypothesis” or “fail to reject the null hypothesis.” A test does not prove an alternative, and a non-significant result does not prove there is no effect. If the goal is to establish that an effect is smaller than a meaningful threshold, plan an equivalence or non-inferiority analysis rather than interpreting a large p-value as proof of no difference.

Quick test-selection guide

Question Common starting point
Is one sample mean different from a fixed benchmark? One-sample t-test
Do two independent groups have different means? Welch’s t-test
Do paired measurements differ? Paired t-test; Wilcoxon signed-rank if a rank-based paired analysis is appropriate
Do three or more independent groups have different means? ANOVA; Welch’s ANOVA if variances differ
Are two categorical variables associated? Chi-square test of independence; Fisher’s exact test may suit a sparse 2×2 table
Are two numeric variables linearly associated? Pearson correlation
Are two numeric or ordinal variables monotonically associated? Spearman correlation; consider Kendall correlation for rank association
Do two independent groups differ in their distributions or ranks? Mann–Whitney U
Are three or more independent groups compared without a normal-error model? Kruskal–Wallis, with suitable follow-up comparisons
Are observations paired, repeated, or clustered? Paired, repeated-measures, cluster-robust, or mixed-effects method appropriate to the design
Were many outcomes, features, or segments tested? Pre-specify primary comparisons or apply a multiplicity correction

Before selecting a test, identify the outcome scale and estimand (mean, proportion, rank effect, or association), whether observations are independent or paired, the number of groups, missingness, and influential observations. Normality is not an automatic pass/fail gate: inspect the relevant residuals or paired differences and consider sample size and design.

1. t-tests: compare means

A t-test assesses a hypothesis about a mean. Use a one-sample test to compare a sample mean with a benchmark, an independent-samples test for unrelated groups, and a paired test for measurements linked within the same unit. For two independent groups, Welch’s t-test is often a safer default when equal population variances are not justified. In SciPy, ttest_ind defaults to equal variances; set equal_var=False for Welch’s test. See the SciPy t-test documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When to use one

  • Compare average revenue per user between treatment and control.
  • Compare mean latency for two infrastructure configurations.
  • Compare average model error under two methods.
  • Compare before-and-after measurements for the same customers using a paired test.

Assumptions to check

  • The outcome is quantitative and observations are independent between groups.
  • For classical mean-based inference, the sampling behavior should be reasonably suited to the t procedure; severe skew, heavy tails, and outliers can matter, especially in small samples.
  • Student’s pooled-variance test assumes equal population variances; Welch’s test does not.
  • For a paired t-test, assess the within-unit differences, not whether each raw group separately looks normal.

Run Welch’s or a paired test in Python

from scipy import stats

# Independent groups: Welch's two-sided test
result = stats.ttest_ind(
    treatment,
    control,
    equal_var=False,
    alternative="two-sided"
)
print(result.statistic, result.pvalue, result.df)
print(result.confidence_interval())

# Paired measurements: test the within-unit differences
paired = stats.ttest_rel(after, before, alternative="two-sided")
print(paired.statistic, paired.pvalue, paired.df)

The independent result reports a statistic, p-value, and degrees of freedom; the returned result object also provides a confidence interval in the documented SciPy API. For a paired comparison, report the mean within-unit difference and its interval. Include group means, the test variant, and an effect size such as Cohen’s d or Hedges’ g.

Common mistakes and alternatives

  • Do not use a pooled Student’s test when equal variances are implausible, especially with unequal group sizes.
  • Do not count repeated rows from one user as independent observations.
  • Do not remove outliers solely to obtain a smaller p-value; investigate their provenance and assess sensitivity.
  • For ordinal or skewed independent samples, consider Mann–Whitney U; for paired data, consider Wilcoxon signed-rank. Permutation tests, robust or trimmed-mean procedures, regression, and generalized linear models may better fit other designs or outcomes.

2. Chi-square tests: categorical counts and association

The chi-square test of independence asks whether two categorical variables are associated. For example, it can test whether conversion status is related to treatment assignment. Goodness-of-fit tests instead compare observed counts with a specified distribution; tests of homogeneity compare categorical distributions across populations.

Requirements and diagnostics

  • Use frequency counts for categorical data, not arbitrary raw continuous values.
  • Observations should be independent.
  • Inspect expected cell counts; a sparse table can make the chi-square approximation unreliable.
  • Examine row and column percentages and observed versus expected counts to understand the pattern, not just its overall test.

Run a test of independence in Python

import pandas as pd
from scipy.stats import chi2_contingency

table = pd.crosstab(df["group"], df["converted"])
chi2, p_value, dof, expected = chi2_contingency(table)
print("chi-square:", chi2)
print("p-value:", p_value)
print("degrees of freedom:", dof)
print("expected counts:n", expected)

See the SciPy contingency-table reference. Describe association strength with an appropriate estimate, such as Cramér’s V; for a 2×2 table, an odds ratio or difference in proportions may be more directly useful. Provide uncertainty intervals where appropriate.

When another method fits better

For a sparse 2×2 table, Fisher’s exact test may be preferable. Other exact procedures can be appropriate in some settings. McNemar’s test is designed for paired binary outcomes. Logistic regression can adjust for covariates; multinomial or ordinal regression may fit more complex categorical outcomes. A significant chi-square result establishes evidence of association under the analysis, not causation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. ANOVA: compare means across groups

One-way ANOVA tests the omnibus null that independent groups share a common mean: H0: μ1 = μ2 = … = μk. A significant result indicates evidence that at least one mean differs; it does not identify which groups differ. ANOVA uses an F statistic to compare variation between groups with residual variation, which is why listing an F-test as a separate everyday test often duplicates the ANOVA use case.

Run one-way ANOVA in Python

from scipy import stats

result = stats.f_oneway(
    df.loc[df["plan"] == "basic", "revenue"],
    df.loc[df["plan"] == "pro", "revenue"],
    df.loc[df["plan"] == "enterprise", "revenue"]
)
print(result.statistic, result.pvalue)

Consult the SciPy one-way ANOVA reference for the function’s current behavior. This example is for a basic independent-groups comparison; factorial, repeated-measures, and covariate-adjusted questions require a model that represents those features.

Assumptions and follow-up comparisons

  • Observations are independent, the outcome is quantitative, and the residuals are reasonably suited to the model, particularly in small samples.
  • Classical ANOVA assumes similar group variances; unequal variances and very different group sizes can be problematic.
  • Severe outliers can dominate mean-based comparisons.
  • After an omnibus result, use planned contrasts or corrected post-hoc comparisons to determine which differences are supported. Tukey’s method is common for all pairwise comparisons; Dunnett’s compares groups with a control; Games–Howell is an option when variances differ.

Do not run every possible uncorrected pairwise t-test after ANOVA. Report group means, a confidence interval for contrasts, and an effect size such as eta-squared or omega-squared. If variances differ, consider Welch’s ANOVA; if observations are repeated or clustered, use a repeated-measures or mixed-effects model. Kruskal–Wallis, permutation approaches, and generalized linear models are alternatives when the outcome or assumptions call for them.

4. Correlation tests: quantify association

Correlation tests assess association according to a particular definition. Pearson correlation measures linear association; Spearman correlation measures rank-based monotonic association. Both coefficients range from −1 to +1, but a value of zero means no association of that specified form, not necessarily no relationship at all.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pearson: linear association

from scipy.stats import pearsonr

result = pearsonr(df["ad_spend"], df["revenue"])
print(result.statistic, result.pvalue)
print(result.confidence_interval())

See the SciPy Pearson correlation reference. Report the coefficient and interval, and inspect a scatterplot. Pearson correlation is sensitive to outliers and can miss curved relationships.

Spearman: monotonic or rank association

from scipy.stats import spearmanr

rho, p_value = spearmanr(df["ranked_feature"], df["ranked_outcome"])
print(rho, p_value)

Spearman is useful for ordinal measurements or relationships that are monotonic but not necessarily linear. Its API is documented in SciPy’s Spearman reference.

Interpret with care

Independence matters for ordinary correlation inference. A high correlation may reflect a confounder, time trend, selection bias, or shared denominator. For time series, account for trends and autocorrelation rather than treating each point as independent. If the real goal is prediction or causal estimation, regression or a causal design may answer it more directly. For broader dependence questions, consider Kendall correlation, partial correlation, nonlinear models, or other methods suited to the data.

5. Mann–Whitney U: compare two independent samples by ranks

The Mann–Whitney U test compares two independent samples through their ranks. Its general null concerns the underlying distributions; it is not universally a test of medians. A location or median interpretation needs additional conditions, including comparable distribution shapes. The SciPy Mann–Whitney documentation describes its hypotheses and calculation methods.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When it can help

  • Compare ordinal satisfaction scores or skewed transaction values across independent groups.
  • Assess rank differences when a mean-based model is not the intended analysis.
  • Use only when observations are independent; it is not the paired-sample procedure.

Run the test in Python

from scipy.stats import mannwhitneyu

result = mannwhitneyu(
    treatment,
    control,
    alternative="two-sided",
    method="auto"
)
print(result.statistic, result.pvalue)

In the documented SciPy API, method="auto" chooses an exact calculation for sufficiently small samples without ties and an asymptotic calculation otherwise. The exact method does not correct for ties; for small tied samples, a permutation method may be more appropriate. Report a rank-based effect estimate, such as rank-biserial correlation or probability of superiority, with uncertainty where feasible.

Limits and alternatives

Calling Mann–Whitney a “nonparametric t-test” can mislead: it may address a different estimand from a mean comparison, and rank-based methods are not assumption-free. If distribution shapes differ substantially, avoid interpreting the result as a simple location shift. For paired samples consider Wilcoxon signed-rank; for concerns about unequal shapes, Brunner–Munzel may suit the question. Quantile regression, robust methods, or a permutation analysis may also be useful.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When a simple test is not enough

Respect dependence and the unit of analysis

Pseudoreplication—treating correlated rows as independent—is a common source of false precision. Multiple sessions may belong to one user; measurements may be repeated on one device; patients may be nested within clinics. Use a paired or repeated-measures method, a mixed-effects model, cluster-robust inference, or a justified aggregation strategy. For time-dependent observations, account for serial correlation.

Check missingness, outliers, and assumptions

Document whether the analysis uses complete cases, imputation, weighting, or another missing-data approach. Missingness related to treatment, outcome, or group can undermine a simple comparison. Investigate outliers and data-quality issues rather than deleting points automatically; compare robust, transformed, trimmed, or permutation analyses as sensitivity checks when appropriate. Normality tests alone are poor gatekeepers: in large samples they can flag negligible departures, while in small samples they may have little power. Use plots, design knowledge, and diagnostics of the relevant residuals or differences.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For experiments, the design supports the conclusion

In a randomized A/B test, Welch’s t-test can compare a continuous metric, while a proportion test, chi-square test, logistic regression, or count model may suit binary or count outcomes. No single test is best for every metric. A causal interpretation depends on valid random assignment, appropriate randomization unit, correct treatment exposure, limited interference, adequate duration and sample size, and careful handling of attrition and novelty effects. A test on observational data alone generally supports an association or group-difference claim under its model, not a causal claim.

Control false positives when testing many hypotheses

When many metrics, segments, features, or pairwise contrasts are tested, the chance of at least one false positive rises. Decide which outcome is primary before examining results where possible, distinguish exploratory findings from confirmatory ones, and choose a correction suited to the decision. Bonferroni and Holm procedures control family-wise error; Benjamini–Hochberg controls the false discovery rate under its conditions. For ANOVA follow-ups, use a method designed for the comparison family, such as Tukey or Dunnett where appropriate. The statsmodels multiple-testing API implements common correction procedures.

Corrections do not repair biased sampling, dependence violations, poor measurement, or a post hoc story. Report how many hypotheses were examined and which correction was applied.

Report results so readers can judge their importance

A useful result statement names the comparison, estimate, uncertainty, procedure, and practical context. For example, report a difference in means with its confidence interval, Welch’s test statistic and degrees of freedom, p-value, and an effect size—not just “significant.” For categorical outcomes, include proportions and an association measure; for ANOVA, include group summaries and follow-up contrasts; for correlations, give the coefficient and interval; for rank tests, explain the rank-based effect rather than translating it automatically into a median difference.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Statistical significance is not practical significance. Compare the estimated effect and its interval with a business or scientific threshold. A narrow interval around a negligible effect may rule out useful impact; a broad interval may leave meaningful effects plausible even when the p-value exceeds the chosen threshold.

A practical workflow for choosing and running a test

  1. Define the question. State the null and alternative in terms of a specific estimand, such as a mean difference, proportion difference, or association.
  2. Identify the outcome and predictors. Decide whether the outcome is numeric, categorical, count, or ordinal, and whether the question concerns groups or association.
  3. Map the design. Establish whether observations are independent, paired, repeated, or clustered; define the unit randomized or sampled.
  4. Choose a method. Use the decision table as a starting point; select a model when covariates, non-Gaussian outcomes, or complex dependence require it.
  5. Inspect data and assumptions. Check missingness, data quality, outliers, residuals or differences, variance patterns, and expected categorical counts.
  6. Plan uncertainty and multiplicity. Choose the effect size, confidence interval, significance threshold, and any multiple-testing correction before interpreting results.
  7. Report the estimate and limits. Give the test variant and result, practical magnitude, uncertainty, design constraints, and a conclusion no stronger than the design supports.

The central rule is to choose from the question and data-generating design, not from a memorized list of test names.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.