Adjust alpha when several hypotheses form one decision-relevant family and you may emphasize or act on findings because their p-values are small. The number of analyses alone is not the trigger: first define the claim and the tests from which you could select results, then choose an error rate and a procedure that fit the consequences of a false positive.
When does multiple testing call for an adjustment?
Multiplicity matters when results from several tests can influence the same scientific claim or decision. If you examine multiple endpoints, subgroups, outcomes, or model specifications and then give special weight to whichever results have the smallest p-values, the selection process makes false-positive findings more likely. Adjusting accounts for that selection.
A useful guiding principle is to adjust when reporting or interpretation places more emphasis on one or more results because their p-values are small. Whether adjustment is warranted is a different question from which method to use.
Define the claim and its family of tests
A family is the set of hypotheses that could be selected interchangeably to support the same decision or claim. It is an inferential choice, not automatically every variable, test, or model in a database. Define it before looking at results wherever possible; otherwise, decisions about which tests count may be influenced by the findings.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitches#1 Best Overall
For example, if a study will support a treatment claim when any one of several clinical endpoints is significant, those endpoints belong in the relevant family. By contrast, tests answering unrelated questions that cannot substitute for one another do not necessarily need to be pooled into one family. State the rationale for the boundary so readers can see what conclusions the procedure covers.
Distinguish confirmatory tests from descriptive or exploratory work
Purely descriptive analyses do not automatically require a correction if no selective inferential decision is being made. But if exploratory results are chosen for emphasis because their p-values are small, present them as exploratory and avoid treating unadjusted p-values as confirmatory evidence. Report analyses added after seeing the data rather than making them appear prespecified.
Rank #2
- New
- Mint Condition
- Dispatch same day for order received before 12 noon
- Guaranteed packaging
- No quibbles returns
Choose the error rate to match the decision
The main choice is between controlling the family-wise error rate (FWER) and controlling the false discovery rate (FDR). They answer different questions, so one is not universally better than the other.
| Target | What it controls | Typical purpose | Main trade-off |
|---|---|---|---|
| FWER | The probability of making one or more false rejections in the family. | Confirmatory decisions where even one false positive could lead to an unacceptable scientific, clinical, regulatory, or product decision. | Can be conservative when many hypotheses are tested, reducing power to detect real effects. |
| FDR | The expected proportion of false discoveries among rejected hypotheses. | Discovery work across many hypotheses where some reported findings may be false, provided the expected false-discovery proportion is controlled. | Does not promise that every discovery, or the family as a whole, is free of false positives. |
Use FWER when one false positive matters greatly
In a clinical trial, a false conclusion about any one of several endpoints may affect how a drug’s effects are judged. The FDA warns that as the number of endpoints analyzed in one trial increases, the chance of false conclusions about one or more endpoints becomes a concern without appropriate multiplicity handling. This is why confirmatory trials commonly specify endpoint hierarchies and multiplicity plans before unblinding.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Use FDR for a discovery list
For large-scale discovery, a team may prefer more power and accept that some reported discoveries could be false. The Benjamini–Hochberg procedure, introduced in 1995, controls FDR under its applicable conditions and can be more powerful than common FWER procedures. Its target is the expected share of false findings among rejections, not the probability of at least one false finding.
Which adjustment procedure should you use?
Choose a method based on the error rate you need, the dependence among tests, any ordering or weighting, and the cost of a false positive. Do not choose a familiar label without specifying what its guarantee covers.
Rank #4
| Procedure | Target | How it works | Considerations |
|---|---|---|---|
| Bonferroni | FWER | For m tests and family alpha α, test each hypothesis at α/m, or multiply each raw p-value by m. | Simple and valid under arbitrary dependence, but can be conservative. |
| Holm step-down | FWER | Sort p-values from smallest to largest. Compare the smallest with α/m, the next with α/(m−1), and continue until a comparison fails; reject only through the last consecutive pass. | Controls FWER under arbitrary dependence and is at least as powerful as unmodified Bonferroni. R’s official documentation says there is generally no reason to use unmodified Bonferroni when Holm is available. |
| Hochberg, Hommel, or Sidak | FWER | Use their procedure-specific adjusted comparisons or p-values. | Validity and power depend on assumptions such as the dependence structure and the inferential objective; justify the choice rather than selecting it by habit. |
| Benjamini–Hochberg (BH) | FDR | Sort p-values p(1) ≤ … ≤ p(m). For target q, find the largest k such that p(k) ≤ kq/m, then reject hypotheses with ranks 1 through k. | Document the family, dependence assumptions, filtering, and any weighting. BH is not an FWER correction. |
| Benjamini–Yekutieli (BY) | FDR | Uses an FDR procedure designed to work under broader dependence conditions than BH. | Usually more conservative than BH. |
Here, m is the number of hypotheses in the declared family; α is the chosen family-wise error target, and q is the chosen FDR target. The Holm and BH formulas above are decision rules, not claims that any particular alpha or q is appropriate for every study.
Use hierarchy, gatekeeping, or weighting when the scientific plan requires it
When endpoints have a prespecified priority or hypotheses should be tested only after earlier conditions are met, a hierarchy or gatekeeping procedure may fit better than treating every test equally. If tests receive different weights or alpha allocations, specify the rules and rationale in advance. These choices define which claims can be supported and how the error-rate guarantee is maintained.
A practical sequence for planning and reporting
- Write the claim. Decide whether the conclusion concerns one designated endpoint, any endpoint, all endpoints, or a list of discoveries.
- List the selectable hypotheses. Include the tests that could be highlighted or acted on as evidence for that claim. Explain why tests outside the family answer a separate question.
- Choose the error-rate target. Use FWER when any false rejection is unacceptable; use FDR when discovery is the goal and a controlled expected proportion of false discoveries is acceptable.
- Specify the method and rules. Record the target level, procedure, ordering, weighting, hierarchy, gatekeeping, and alpha allocation. In clinical trials, define the endpoint hierarchy and multiplicity handling before unblinding.
- Report what was done. Name the family, number of tests, target, method, and whether you adjusted p-values or used adjusted thresholds. Give raw and adjusted p-values, or the exact thresholds used, as appropriate.
- Explain implications for intervals and conclusions. If confidence intervals are used to support multiplicity-adjusted claims, state whether and how they were adjusted. Separate confirmatory conclusions from exploratory findings, and disclose analyses added after viewing the data.
Common mistakes that change the meaning of a result
- Correcting across unrelated analyses without a decision rationale: pooling tests that cannot substitute for one another may unnecessarily reduce power and obscure which claim the adjustment protects.
- Ignoring selection after searching: testing many endpoints, subgroups, outcomes, or model specifications and then reporting only the smallest p-values does not become confirmatory because the final report lists just a few results.
- Calling BH an FWER method: BH targets FDR under its conditions; it does not control the probability of any false rejection in the family.
- Reporting only “Bonferroni corrected”: without the family, test count, target, and whether p-values or thresholds changed, readers cannot reconstruct the analysis.
- Equating adjusted significance with importance: an adjusted p-value does not show that an effect is large, precise enough for a practical decision, or worth acting on. Interpret effect size, uncertainty, and consequences separately.
What an adjustment can—and cannot—support
A multiplicity procedure supports a specified set of inferences under its stated assumptions. It cannot repair an unclear family definition, selective reporting, an unsuitable model, or a result that lacks practical importance. Make the decision question, family, and error-rate target visible so readers can judge what the reported findings establish.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

