Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Correlation shows that two things vary together; it does not, by itself, show that one causes the other. The association may reflect cause and effect, reverse causation, a shared cause, chance, or bias. To judge causation, define the question, examine how the data were produced, and weigh the result alongside assumptions and other evidence.

What correlation tells you—and what it leaves open

Correlation is a pattern of association: values of two variables tend to change together in a dataset or population. The pattern is useful evidence that may point toward a relationship worth investigating, but it does not identify the explanation. The same observed correlation can fit several different causal stories, as Harvard’s overview of causal and non-causal associations explains.

  • X causes Y: changing X would change Y, all else relevant held appropriately constant.
  • Y causes X: the outcome influences the proposed cause; this is reverse causation.
  • A third factor affects both: a confounder helps create the observed association.
  • The pattern is misleading: chance, selection, measurement, or other study problems can produce or distort an association.

For example, ice-cream sales and drowning incidents could both increase during warm weather. Temperature is a plausible common cause in this hypothetical; the example is not a measured finding about a particular place or period. Likewise, if illness is associated with a behavior, illness might have changed the behavior rather than the behavior having caused the illness. The observed association alone cannot settle the direction.

Why a correlation can mislead

Confounding

A confounder is a factor that influences both the exposure being studied and the outcome. If it is not adequately addressed, the exposure can appear related to the outcome even when the association is partly or wholly explained by that shared cause. Which factors matter depends on the causal question and how the data were generated.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reverse causation and timing

To claim that X causes Y, X must occur before the relevant change in Y. A single cross-sectional snapshot often cannot establish that order: the outcome may affect the exposure, or both directions may operate. Repeated measurements or a design that establishes timing can help, but timing alone does not eliminate other explanations.

Chance, selection, and measurement

An apparent association can arise by chance, or because the people included in a study differ systematically from those left out. Inaccurate or inconsistent measurement can also create or distort a pattern. The CDC Field Epidemiology Manual cautions that an observed association may represent a causal connection, but may instead result from chance, selection bias, information bias, confounding, or other errors in design, execution, or analysis (“Analyzing and Interpreting Data”).

A small p-value, statistical significance, or a large correlation does not rule out these alternatives. Such statistics describe patterns under particular analytical assumptions; they are not, on their own, tests that establish causation.

How researchers make a causal question answerable

“Does X cause Y?” can be too vague to guide a study. A sharper question specifies what exposure or intervention is being compared, the alternative, the target population, the outcome, and the time period. For example: among a defined group, what would happen to an outcome over a stated period under one option compared with another?

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The formal potential-outcomes approach asks us to compare what would happen to the same target population under each option. For any one person, however, we cannot observe both alternatives at once. One is observed; the other is a counterfactual. That missing comparison is why causal conclusions depend on study design and assumptions, not just on a statistical association. Hernán and Robins develop this framework in Causal Inference: What If.

What different study designs can establish

Design How exposure is assigned What it can address Key limitations
Randomized experiment Participants or units are assigned by chance to groups. Random assignment helps balance alternative explanations between groups on average, making this a strong design for estimating an intervention’s effects when well conducted. It may be impractical or unethical. Attrition, noncompliance, poor measurement, and limited similarity between study participants and other populations can affect interpretation.
Observational study People or circumstances are not randomly assigned to the exposure. It can reveal patterns and contribute causal evidence when the question, causal model, measurement, and analysis are carefully considered. Groups may differ in measured or unmeasured ways. Adjustment cannot automatically remove unmeasured confounding, and results depend on assumptions.
Natural or quasi-experiment An external change creates differences in exposure, timing, or eligibility that may approximate random assignment. It can support causal inference when the comparison is plausibly as-if random and the design’s assumptions are credible. The label does not guarantee a fair comparison. The reason groups differ and the assumptions needed to interpret the result must be defended.

The U.S. National Library of Medicine describes randomized controlled trials as among the designs most likely to determine a causal relationship, while recognizing that experiments may not be feasible or ethical (“Finding and Using Health Statistics”). Randomization is powerful because assignment is not chosen in response to participants’ characteristics. It does not make every later problem disappear: what happens after assignment, how outcomes are measured, and who the results apply to still matter. Guidance from NICHD likewise distinguishes manipulation and random assignment from the adjustments used in correlational studies when experiments are unavailable (Using Research and Reason in Education).

How to assess a causal claim in practice

  1. Specify the comparison. Identify the exposure or intervention, the alternative, the population, the outcome, and the time window. Be clear about what counts as each option.
  2. Check whether exposure came before outcome. Look at when each was measured and whether reverse causation remains plausible.
  3. Understand how groups were formed. In a randomized experiment, check whether the randomized comparison was preserved. In an observational study, ask why people received or encountered different exposures. For a quasi-experiment, ask why the external change makes the groups plausibly comparable.
  4. Identify likely alternative explanations. State a causal model before choosing adjustment variables. Consider common causes, selection into the study, and how exposure and outcome were measured.
  5. Inspect the assumptions and uncertainty. Regression, matching, and other adjustments can address measured factors only under suitable assumptions; they do not automatically fix a weak design or account for unmeasured factors. Report uncertainty and relevant limitations.
  6. Look for converging evidence. Consider whether the proposed mechanism is plausible, whether findings recur across designs or populations, whether a dose-response pattern is appropriate, and whether negative controls, falsification checks, or reasonable alternative analyses change the result.

These checks strengthen or weaken a causal case; none is a magic proof. The National Academies’ Reference Guide on Statistics and Research Methods emphasizes that observational and quasi-experimental conclusions require attention to assumptions and scientific judgment. A dose-response pattern can add weight, as the CDC notes, but it does not by itself exclude confounding or bias.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Can correlation ever be evidence of causation?

Yes. An association can be one piece of evidence in a causal argument, especially when it is temporally ordered, measured credibly, consistent with a plausible mechanism, and supported by a design that addresses competing explanations. But correlation alone cannot distinguish cause from reverse causation, confounding, chance, selection, or measurement problems. As the NIST/SEMATECH Engineering Statistics Handbook puts it, correlation by itself does not establish a causal relationship.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.