P-Hacking
Trying enough analyses on the same data until one crosses the significance threshold, then reporting only that one.
P-hacking is the practice of running many possible analyses on a single dataset — different subgroups, different variables, different cutoffs, different statistical tests — and reporting only the one that happens to cross the conventional threshold for statistical significance (typically p < 0.05), without disclosing how many other analyses were tried and discarded. The name plays on the p-value itself, the statistic being manipulated, and the practice was named and quantified in an influential 2011 paper by Simmons, Nelson, and Simonsohn, who showed that a few common, individually defensible research choices, stacked together, can push a false effect's apparent significance rate above 60 percent.
The mechanism exploits how a p-value is actually defined: it's the probability of seeing a result this extreme if there were truly no effect and if this were the only test run. Every additional way a researcher is free to slice the data — stop collecting early once the result looks good, drop an inconvenient outlier, try three related outcome measures and report the one that worked, split by age or gender post hoc — adds another independent roll of the dice, and with enough rolls, a random fluctuation will eventually clear the bar even in data generated by pure noise. None of these individual choices needs to be made in bad faith; each can look, from inside a single analysis, like an ordinary judgment call, which is what makes the practice so much more common than outright fraud and so much harder to police from the outside.
The standard defenses are structural, matching the mechanism directly: preregistering the hypothesis and the exact analysis plan before seeing the data removes the researcher's freedom to choose after the fact, and reporting every test actually run, not just the significant one, restores the reader's ability to judge how much the result should be discounted. P-hacking and Publication Bias compound each other: a p-hacked result is more likely to look novel and significant, which is exactly the profile publication bias rewards.
See also4
Publication Bias
Studies that find a significant, interesting effect get published far more often than studies that find nothing.
Meaning & Society4 connections
Selection Bias
A sample distorted because inclusion in it was never random.
Method16 connections
Placebo Effect
An inert treatment produces a real improvement because of the expectation and ritual around it, not any active ingredient.
Meaning & Society7 connections
Availability Cascade
Repetition of a claim raises its perceived truth and prominence, and the perceived prominence drives further repetition.
Meaning & Society13 connections