Base Rate Fallacy
judging a specific case from vivid evidence while discounting how common the categories actually are.
The base rate fallacy is the error of estimating a probability from specific, vivid evidence while ignoring the known frequency — the base rate — of the categories involved. Daniel Kahneman and Amos Tversky demonstrated it with the taxicab problem: told that 85% of a city's cabs are Green and 15% Blue, and that a witness identifies the cab in an accident as Blue with 80% reliability, most people answer that the cab was probably Blue. Worked through Bayes's rule the correct answer is closer to 41%; the rarity of Blue cabs matters as much as the witness's accuracy, and most people give it almost no weight.
Their lawyer-engineer problem shows the same gap. Participants read a sketch of Jack — 45, married with four children, conservative, careful and ambitious, no interest in political or social issues, hobbies of home carpentry, sailing and mathematical puzzles — drawn from a pool stated to be mostly engineers or mostly lawyers, and their guesses barely moved with the stated ratio. (The shy, tidy, people-averse profile often attached to this experiment belongs to a different one, Steve the librarian.) The sketch feels informative; the ratio feels like a technicality attached to the question rather than a fact about the answer.
The mechanism is representativeness: a description resembling a stereotype is judged probable in proportion to how well it resembles it, not to how often that category occurs. Frequency data feels like it belongs to a different question than the one in front of you; a specific case feels self-contained. This is a claim about reasoning, not about instruments — for how a test's own error rates interact with a population's base rate, see Sensitivity and Specificity.
The fallacy recurs anywhere a rare-category judgment is made from a compelling profile: forecasting, screening, and Automation Bias toward a system's specific output over its known error rate. The Monty Hall Problem is a neighbor rather than an instance: it turns not on a neglected base rate but on failing to condition on the host's constrained choice of door, which carries information. Stating the base rate explicitly, as a habit rather than an afterthought, is most of the fix.
See also5
Sensitivity and Specificity
The two error rates of a test, and why neither answers the question a person actually asks.
Method19 connections
Fermi Estimation
Reaching a defensible order-of-magnitude answer by decomposing a question into estimable factors.
Method23 connections
Anchoring Effect
A stated number pulls subsequent estimates toward it, regardless of relevance.
Method35 connections
Selection Bias
A sample distorted because inclusion in it was never random.
Method16 connections
Streetlight Effect
Looking where the light is good rather than where the answer is.
Method12 connections
Related3
Nearby in the graph rather than deliberately chosen. Looser, sometimes surprising.
Linked from10
- Anchoring EffectMethod
A stated number pulls subsequent estimates toward it, regardless of relevance.
- Availability CascadeMeaning & Society
Repetition of a claim raises its perceived truth and prominence, and the perceived prominence drives further repetition.
- Availability HeuristicMethod
Judging how likely or common something is by how easily examples come to mind, not by actual frequency.
- Confirmation BiasMethod
The tendency to search for, interpret, and recall evidence in ways that favor what you already believe.
- Correlation and CausationMethod
Two things moving together is evidence for a causal link but never proof of one, or of its direction.
- Dunning-Kruger EffectMethod
Low skill removes the very ability needed to recognize low skill, inflating self-assessment most at the bottom.
- Hofstadter's LawMethod
A task always takes longer than expected, even when you plan for it to take longer than expected.
- Regression to the MeanMethod
An extreme measurement tends to be followed by a more average one, with no cause required beyond noise.
- Sensitivity and SpecificityMethod
The two error rates of a test, and why neither answers the question a person actually asks.
- Simpson's ParadoxMethod
A trend appears in several groups of data but reverses or disappears when the groups are combined.