xkcd’s “Significant” comic shows why finding one result with p < 0.05 after testing 20 jelly-bean colors is not enough to conclude that green jelly beans cause acne. The green result could be real, but searching across many comparisons makes a chance finding more likely—and reporting only the standout result hides how it was found.
What happens in xkcd’s “Significant” comic?
In xkcd comic 882, the researchers first test whether jelly beans generally are linked to acne and find no link. They then test the colors separately. Green is the one color shown with p < 0.05; the other color comparisons are shown with p > 0.05. A newspaper turns that result into “Green Jelly Beans Linked To Acne!” and “95% Confidence.”
The comic’s title is “Significant.” Its alt text supplies the second punchline: “So, uh, we did the green study again and got no link,” followed by the newspaper’s claim that research is conflicted and more study is recommended. The comic is a satire about how an unconfirmed finding can be framed, not a report of a real acne experiment.
Why testing 20 colors changes the interpretation
A p-value threshold applies to an individual test. If researchers run many tests, there are more chances for at least one to cross that threshold by chance. The Springer Nature chapter “A Reckless Guide to P-values,” section 3.2, puts the general warning plainly: “The more hypothesis tests there are, the higher the risk that one of them will yield a false positive result.”
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
For 20 independent tests where every null hypothesis is true, using a 0.05 threshold for each test gives an expected one false positive per 20 tests on average. That expectation is not the same as saying the chance of at least one false positive is 5%; the 5% is the per-test threshold, not the error rate for the whole search.
How a Bonferroni correction would work here
One simple way to control the family-wise false-positive rate—the chance of at least one false positive across a set of tests—is the Bonferroni correction. For 20 tests and a 5% family-wise target, divide 0.05 by 20: each test would need to meet a threshold of 0.0025. This is an illustrative calculation, not a rule that every analysis must use Bonferroni. It can reduce statistical power, and another method or a prespecified analysis plan may be more suitable.
Rank #2
- This guide is a perfect overview for the topics covered in introductory statistics courses.
The comic does not provide the exact green p-value, sample size, study design, or data. Because the value is only described as below 0.05, the comic alone cannot show whether the result would meet the 0.0025 threshold.
Does p < 0.05 mean there is a 95% chance the result is true?
No. A p-value below 0.05 does not mean there is a 95% probability that the hypothesis is true, nor does it mean there is only a 5% chance that the result is coincidence. It describes how the observed data, or more extreme data, compare with a specified null model. It does not by itself give the probability that the claim is correct.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #3
The newspaper’s “95% Confidence” headline captures the misunderstanding: a cutoff used for an individual test is presented as certainty about the broad claim. That is not what the comic’s p-value establishes.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Is the green result necessarily false?
No. Multiple testing makes a selected result less conclusive than it might appear when shown alone, but it does not prove that the association is false. A green finding could be a useful lead for a new hypothesis. The important distinction is between exploratory evidence that suggests what to study next and confirmatory evidence that tests the claim using a plan set in advance.
Quick Recap
Best Value
Rank #4
How should researchers report a result found this way?
- Disclose the search. Report that 20 colors were tested, rather than presenting green as if it were the only comparison made.
- Describe the finding as exploratory. A selected result can motivate follow-up, but the selection process belongs in the interpretation.
- Choose an appropriate analysis plan. Decide how to handle multiple comparisons based on the study’s goals; Bonferroni is one option, with a potential cost in power.
- Test the hypothesis with independent data. Evaluate the green association in a fresh study rather than treating the original selected result as confirmation.
- Report both rounds fully. Include the initial comparisons and the follow-up result, including the comic’s kind of outcome: a repeat study that finds no link.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

