Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A results section should report each important estimate together with its confidence interval, usually at the 95% level, and place the p-value after them rather than in their place. The interval shows how large an effect appears to be and which values remain compatible with the data. A p-value alone cannot show that. Intervals do not, however, correct a biased design, an invalid model, or selective reporting.

Report the estimate, then the interval, then the p-value

The American Heart Association and American Stroke Association (AHA/ASA) author guidance on statistical reporting sets out an explicit order: “Quantitative results should be presented in the following order: the estimated effect size (point estimate), the confidence interval (typically 95%) followed by the associated actual p-value.” The JAMA Network instructions for authors take a similar position. They ask authors to quantify findings with appropriate uncertainty indicators, such as confidence intervals, and they warn against relying solely on hypothesis testing.

Writing a result sentence that follows this order takes five decisions:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Name the contrast and the reference group, such as “treatment compared with placebo” or “women compared with men.”
  2. Give the point estimate with its units.
  3. Give the confidence interval in brackets, with its level stated if it is not 95%.
  4. Add the actual p-value when it is relevant to the question.
  5. Name the analysis population and the model or method in the text or the table note.

The general pattern is shown in the code block below. The numbers in the worked example that follows are hypothetical and are included only to show the format.

#1 Best Overall
Estimated [effect measure] was [point estimate] (95% CI [lower, upper]; p = [actual value])

A hypothetical example reads: “In the adjusted model, mean systolic blood pressure was 4.2 mmHg lower in the intervention group than in the control group (difference in means, 95% CI 1.1 to 7.3; p = 0.008).” A reader can see the direction, the size, and the plausible range in one sentence.

What the interval does and does not tell you

An interval describes the precision of the estimate and the uncertainty around it. It is not a statement of certainty that the true value lies inside the reported range. The American Physiological Society (APS) guidance, first issued in 2004, explains the idea through repeated sampling. If the same procedure were applied to many hypothetical samples, the stated proportion of the intervals it produced would contain the fixed population value. The guidance uses 200 hypothetical samples as an explanatory example. It is not an empirical finding.

This distinction has a practical consequence. A 95% interval comes from a procedure with 95% long-run coverage under its assumptions. It does not mean that a particular computed interval has a 95% probability of containing the parameter. The 2007 APS follow-up puts the purpose plainly: “A confidence interval focuses attention on the magnitude and uncertainty of an experimental result.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
Statistics Laminate Reference Chart: Parameters, Variables, Intervals, Proportions (Quickstudy: Academic )
  • This guide is a perfect overview for the topics covered in introductory statistics courses.

Read the width, not only whether the interval crosses the null

A narrower interval generally indicates a more precise estimate. A wide interval means the data leave many values plausible. Those values may include a worthwhile benefit, no effect, or harm. A wide interval therefore describes imprecision. It is not evidence that the effect is absent.

Interpret the width against two yardsticks: the scale of the outcome and the thresholds that matter for decisions. A 2 mmHg interval may be clinically irrelevant for one question and important for another. Ask whether the lower and upper bounds would lead a clinician, policymaker, or engineer to different conclusions. If they would, the paper should say so.

Know the null value for each effect measure

The null value is the value that represents no effect. Readers need it to judge whether an interval is compatible with no difference. Name the null in the text or the table note.

Rank #3
Effect measure Typical use Null value How to read the interval
Difference in means Continuous outcomes, such as blood pressure or test scores 0 An interval that includes 0 is compatible with no difference between groups.
Risk difference or percentage-point difference Absolute differences in proportions, such as survey percentages 0 An interval that includes 0 is compatible with no absolute difference.
Ratio (risk ratio, odds ratio, hazard ratio) Relative comparisons 1 An interval that includes 1 is compatible with no relative difference.

Describe nonsignificant comparisons accurately

The U.S. Census Bureau’s Statistical Quality Standard E2, Reporting Results, requires uncertainty measures for key estimates. These may be confidence intervals, margins of error, or equivalents in the information products it specifies. It also requires direct comparisons that are not statistically significant to be identified as such. Those requirements are agency conventions. They are not universal rules for every journal or field.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Census standard sets a 90% confidence level for Census Bureau publications and news releases, and 90% or more for other listed information products. Other fields commonly use 95%, so the level should always be stated.

When an interval includes the null value, the correct conclusion is that the data are compatible with a range of effects, including zero. It is not that the groups are equal. A sentence such as “the difference was not statistically significant (difference 1.8 points, 95% CI −0.9 to 4.5)” gives the reader more than “there was no difference.”

Do not turn a threshold into the story

AHA/ASA guidance cautions against drawing conclusions only from whether a p-value passes a threshold. Authors should explain the size of the effect, the uncertainty around it, and its clinical or biological relevance. A result can cross the threshold and still be too small to matter. Another can miss it and still leave a potentially important effect plausible. The interval makes both situations visible.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What intervals cannot fix

A confidence interval describes the sampling and model uncertainty that the analysis captures. It does not address every problem in a study. The following issues remain the author’s responsibility:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Bias and measurement error. An interval around a biased estimate is a precise measurement of the wrong quantity.
  • Confounding and model misspecification. Adjustment choices shape the estimate and its interval.
  • Missing data. How missingness was handled changes what the interval means.
  • Multiple outcomes and comparisons. Many intervals do not correct for multiplicity or selective reporting.
  • Complex designs. Clustered, stepped, or adaptive designs may need an interval method matched to their structure.

The APS guidance itself warns that reporting rules cannot replace an understanding of statistical concepts and procedures. The ARRIVE guidelines, which cover reporting of animal research, also place effect sizes with their interval estimates in the Results section, alongside precision and evidence synthesis. Authors should state the analysis plan and the method used to obtain each interval, especially for complex designs.

When the framework is not frequentist

Bayesian analyses produce credible intervals, not confidence intervals. A credible interval is built and interpreted within a Bayesian framework, and it should be named as such. Its construction and interpretation should be stated, rather than assuming it is interchangeable with a frequentist interval. The same care applies to any other uncertainty summary. Report the appropriate interval for the method actually used.

Readers who want more depth can consult Statistics with Confidence: Confidence Intervals and Statistical Guidelines, second edition, edited by Douglas Altman, David Machin, Trevor Bryant, and Martin Gardner. It includes worked examples and reporting guidance.