Effect size quantifies how large a difference, association, or model contribution is. A p-value helps assess how compatible data are with a null model; it does not say whether an observed effect is practically large. In Python, the right measure depends on the outcome and study design: use a standardized mean difference for continuous outcomes in two groups, a correlation for association, an eta-squared variant for ANOVA, or a measure such as an odds ratio for binary outcomes.
Choose an effect size that matches the question
Before calculating anything, identify the outcome type and whether observations are independent, paired, or repeated. Then choose a measure whose scale and interpretation fit the audience. These measures are not interchangeable, and their numeric values should not be compared as if they shared one scale.
| Question or analysis | Useful measure | What it describes |
|---|---|---|
| How far apart are two group means on a continuous outcome? | Cohen’s d or Hedges’ g | Mean difference expressed in standard-deviation units. |
| How strongly are two variables associated? | Correlation, such as r | Direction and strength of a relationship on a correlation scale. |
| How much variation is associated with an ANOVA effect? | Eta-squared or partial eta-squared | A variance proportion, with the exact variant determining the denominator. |
| How does a binary outcome’s odds differ between groups? | Odds ratio | A multiplicative comparison of odds. |
| How often would a value from one group exceed a value from another? | AUC or common-language effect size | A probabilistic interpretation of group separation. |
Calculate Cohen’s d and Hedges’ g for two independent groups
Cohen’s d uses a pooled standard deviation
For two independent groups, Pingouin documents pooled-standard-deviation Cohen’s d as the difference between the group means divided by their pooled standard deviation:
d = (mean1 − mean2) / sqrt(((n1 − 1)s1² + (n2 − 1)s2²) / (n1 + n2 − 2))
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute#1 Best Overall
The sign depends on the order of the groups: with this convention, a positive value means group 1’s mean is higher than group 2’s. Keep that order explicit when reporting results. The pooled denominator is appropriate to this independent-groups definition; do not apply it silently to paired observations.
Hedges’ g adjusts for small-sample bias
Hedges’ g applies a small-sample correction to Cohen’s d. Pingouin gives the correction as g = d × (1 − 3 / (4(n1 + n2) − 9)). Its documentation warns that Cohen’s d is biased as an estimate of the population effect size, especially for small samples (n < 20). Treat that as the package’s warning threshold, not a universal boundary at which d becomes unusable.
Rank #2
- This guide is a perfect overview for the topics covered in introductory statistics courses.
Compute both with Pingouin
Pingouin is an open-source Python statistical package based mostly on Pandas and NumPy. After installing it in your environment and preparing two numeric collections, compute the estimates and a confidence interval like this:
import pingouin as pg
d = pg.compute_effsize(group_a, group_b, paired=False, eftype="cohen")
g = pg.compute_effsize(group_a, group_b, paired=False, eftype="hedges")
ci = pg.compute_esci(stat=d, nx=len(group_a), ny=len(group_b), eftype="cohen")
Here group_a and group_b are the observations for the two independent groups. The confidence-interval call uses Cohen’s d; if you report Hedges’ g, ensure the interval you present corresponds to the effect-size statistic and method actually used.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
Use a paired definition for matched or repeated observations
Paired designs have more than one defensible standardized-mean-difference denominator. Pingouin documents d-avg, which uses the average of the two variances, and d-z, which uses the standard deviation of the difference scores. They answer slightly different reporting questions, so state the variant rather than labeling either one simply “Cohen’s d.”
For matched or repeated observations, use paired=True in the relevant Pingouin calculation, and report the paired variant or denominator used. Also document how missing values were handled: removing incomplete pairs can change the effective sample size and which observations contribute to the estimate.
Rank #4
Report the exact eta-squared variant for ANOVA
Eta-squared is a variance-proportion measure. Partial eta-squared conditions the proportion on the model’s error and other terms, so it can differ from standard eta-squared. Pingouin’s ANOVA output labels partial eta-squared as np2 and discusses standard eta-squared as an alternative. Preserve the label shown by the analysis and name the variant in the report; do not silently substitute one for the other.
Use measures that fit association and binary outcomes
Correlation and probabilistic superiority
A correlation such as point-biserial r describes association. For a two-group comparison, AUC and common-language effect size express separation probabilistically: the common-language value is P(X > Y) + 0.5P(X = Y), the probability that a value from one group exceeds a value from the other, with ties counted as half.
Best Value
Pingouin documents conversions including d = 2r / sqrt(1 − r²) and AUC = Φ(d/√2). Such conversions can help translate between scales, but each measure answers a different interpretive question. Prefer a measure computed directly for the observed design when that is available, and do not treat converted numbers as independent evidence.
Odds ratios for binary outcomes
An odds ratio compares the odds of a binary outcome multiplicatively. Pingouin lists odds ratio among its supported effect-size types. It also documents OR = exp(dπ/√3) as a conversion from Cohen’s d; treat this as a model-based approximation, not a substitute for calculating an odds ratio directly from the observed binary data when that is the question of interest.
Report the estimate with enough context to interpret it
An effect-size number without its design and uncertainty can be misleading. A useful report identifies the groups and outcome, the direction convention, the measure and denominator or variant, sample sizes, and a confidence interval. Include assumptions or data-handling choices that affect the calculation, such as pairing and missing-value handling.
- For two independent continuous groups, state whether the estimate is Cohen’s d or Hedges’ g and identify group order.
- For paired data, name the paired variant, such as d-avg or d-z.
- For ANOVA, say whether the result is eta-squared or partial eta-squared.
- For binary outcomes, distinguish a directly calculated odds ratio from a conversion based on another effect-size scale.
- Give the confidence interval for the reported effect size, not merely for a different statistic.
Avoid assigning “small,” “medium,” or “large” labels without context. Practical importance depends on the outcome, field, and consequences of the difference; the measure’s magnitude alone does not establish it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

