Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

No. A non-significant result in a replication does not, by itself, make the replication a success, and it does not show that an earlier reported effect is absent. A p-value that fails to cross the chosen threshold tells you that the test did not reach that threshold. It does not tell you whether the effect is zero, whether it is too small to matter, or whether the study was never capable of detecting it. The eLife authors put the core problem directly in their 2024 article, “Replication of null results: Absence of evidence or evidence of absence?”: “Non-significance in both studies does not ensure that the studies provide evidence for the absence of an effect and ‘replication success’ can virtually always be achieved if the sample sizes are small enough.”

Start with the claim, not the verdict

Before you decide what a replication means, reconstruct what it was trying to test. Three things need to line up: the claim itself, the effect size the study was built to detect, and the estimate the study actually produced, with its uncertainty.

  • The claim. Does the replication test the same hypothesis? A different population, setting, dose, exposure window, or outcome measure can turn a copy of the original procedure into a test of a different question.
  • The design effect size. A study planned around a large effect is poorly placed to detect a small one. Look for a power analysis or sample-size justification in the methods. If none is reported, the study’s sensitivity is unknown, and you cannot tell what a null would have ruled out.
  • The estimate and its interval. A point estimate close to zero with a narrow interval supports a very different conclusion from a point estimate close to zero with an interval that spans both zero and effects of real importance. A significance label hides that difference.

Why a non-significant p-value is weak evidence of absence

A significance test asks how surprising the data would be if there were no effect. A non-significant result means the data were not surprising enough to reject that assumption at the chosen threshold. That is a narrower statement than “the effect is zero,” and a narrower one again than “any effect that exists is too small to matter.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sample size is the main driver of the gap. In a small study, the confidence interval around the estimate is wide, and it usually spans both zero and effects that would matter in practice. Such a study will often return a non-significant result whatever the true effect is. This is why the eLife quotation above is more than a technicality: if both the original and the replication are small enough, “both non-significant” can be produced almost automatically. Counting that outcome as replication success does not control the error rates that matter, and it can label inconclusive studies as successful.

#1 Best Overall

Five checks before you count a null

  1. Confirm the replication tested the same claim. Compare the population, the manipulation or exposure, the outcome measure, and the analysis plan with the original. The PLOS Biology article “What is replication?” treats fidelity to the original design and the relevance of the claim as central to interpreting any outcome. A faithful replication of the procedure can still be an unfaithful test of the claim if the measurement changed meaning.
  2. Find the sensitivity of the design. Ask what effect size the study could reliably distinguish from zero. If the design was planned to detect only a large effect, a null is evidence against a large effect at most.
  3. Check whether a smallest meaningful effect was set in advance. A null only supports “no meaningful effect” if the authors defined what counts as meaningful before seeing the data, and gave a reason for that threshold.
  4. Compare estimates, not just labels. Check whether the original estimate sits inside the replication’s interval, whether the replication estimate sits inside the original interval, and how the two compare with a prediction interval. The next section explains why these checks can disagree.
  5. Check whether the method evaluates absence or only tests a point null. A test of the hypothesis “the effect is exactly zero” is not the same as a test of “the effect is smaller than the threshold that matters.”

How the eLife authors compared original and replication results

The eLife article examined replications of original null results and assessed four different criteria. Each asks a different question, so the counts differ. The table reports the figures for the 15 replications in that examined set.

Criterion What it asks Result in the examined set (eLife, 2024)
Original effect estimate inside the replication’s 95% confidence interval Is the original estimate compatible with the replication’s own uncertainty range? 11 of 15 (73%)
Replication effect estimate inside the original’s 95% confidence interval Is the replication estimate compatible with the original’s uncertainty range? 12 of 15 (80%)
Replication estimate inside the 95% prediction interval based on the original Does the replication fall within the range a future estimate would be expected to occupy under the original’s model? 12 of 15 (80%)
Combined meta-analysis of original and replication estimates non-significant Does the pooled estimate fail to reach significance? 10 of 15 (67%)

These figures describe one set of 15 replications and particular criteria. They are not a general replication rate, and they should not be merged into a single success percentage. A non-significant pooled estimate also does not, on its own, quantify evidence for absence.

Rank #2
Sale
Statistics Laminate Reference Chart: Parameters, Variables, Intervals, Proportions (Quickstudy: Academic )
  • This guide is a perfect overview for the topics covered in introductory statistics courses.

Methods for evidence of absence

If the question is whether an effect is small enough to be ignored, a standard significance test is the wrong instrument. Two approaches are built for that question, and each needs its assumptions stated openly.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Equivalence testing

Equivalence testing starts by defining a smallest effect size of interest, or equivalence bounds, before the data are analysed. The question becomes whether the estimate and its uncertainty lie entirely within those bounds. For example, if a team decides in advance that a standardised difference smaller than 0.2 is too small to matter, the replication’s confidence interval would need to sit inside the range from -0.2 to 0.2 to support practical equivalence. This is an illustrative threshold, not a standard one for any field. Evidence for equivalence is only as strong as the justification for the bounds, and a wide interval that crosses a bound cannot support the claim.

Rank #3

Bayes factors

A Bayes factor compares how well the data support one specified hypothesis over another, such as a model with no effect against a model with a specified effect. The result depends on the hypotheses and on the prior distributions and model choices behind them. A Bayes factor is not assumption-free proof that an effect is exactly zero. Reporting how the conclusion changes under reasonable alternative priors shows how much the answer depends on those choices.

Confidence intervals, prediction intervals, and pooled estimates

A confidence interval describes uncertainty around an estimated effect under the model used. A prediction interval describes the range in which a future study’s estimate may fall, given the model and the original estimate. These answer different questions and should not be treated as interchangeable. A pooled estimate combines the two studies into one aggregate, which can be useful, but it inherits the same limits: a pooled result that is not significant may still be too imprecise to rule out meaningful effects.

Reporting and analysis choices

The US Office of Research Integrity describes two concerns that bear directly on null results: selective reporting of results, including null or inconclusive ones, and selecting among multiple analyses the one that best fits a hypothesized result. The Office of Research Integrity guidance on selective reporting of results sets out these concerns. That guidance does not establish how common these practices are in replication studies, so treat them as questions to ask rather than as evidence that a given study did them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Was the analysis plan stated before the data were collected or analysed?
  • Are other analyses or outcome measures reported, and were they reported in the same level of detail?
  • Is the set of replications complete, or were inconclusive studies left out of the summary?
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Reading a null replication in practice

A non-significant replication can point in several directions. The pattern of the interval and the design determines which reading is justified.

Pattern What it supports What to check
Wide interval spanning zero and meaningful effects Inconclusive. The study could not distinguish between the original effect and no effect. Sample size, design sensitivity, and whether the estimate is precise enough for any conclusion.
Narrow interval that excludes the original estimate but includes small effects The effect may be smaller than originally estimated. Whether the original estimate was inflated by a small sample or selective reporting, and whether the smaller effect matters in practice.
Interval inside pre-specified equivalence bounds Practical equivalence at the bounds chosen, if those bounds were justified in advance. Why the bounds were chosen, and whether the test assumptions hold.
Faithful procedure, different population or setting A possible boundary condition: the effect may hold only under a narrower set of conditions. Differences in population, setting, treatment, or outcome measure between the two studies.

A null replication does not prove the original claim false, and a significant replication does not prove it true. Both verdicts depend on what was tested, how precisely it was estimated, and whether the inferential method addresses the question the reader actually cares about.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.