Kelly Mears

Base Rate Fallacy

judging a specific case from vivid evidence while discounting how common the categories actually are.

Method2 min read361 words10 out · 2 in
also calledBase rate neglectPrior neglect

The base rate fallacy is the error of estimating a probability from specific, vivid evidence while ignoring the known frequency — the base rate — of the categories involved. Daniel Kahneman and Amos Tversky demonstrated it with the taxicab problem: told that 85% of a city's cabs are Green and 15% Blue, and that a witness identifies the cab in an accident as Blue with 80% reliability, most people answer that the cab was probably Blue. Worked through Bayes's rule the correct answer is closer to 41%; the rarity of Blue cabs matters as much as the witness's accuracy, and most people give it almost no weight.

Their lawyer-engineer problem shows the same gap. Participants read a sketch of Jack — 45, married with four children, conservative, careful and ambitious, no interest in political or social issues, hobbies of home carpentry, sailing and mathematical puzzles — drawn from a pool stated to be mostly engineers or mostly lawyers, and their guesses barely moved with the stated ratio. (The shy, tidy, people-averse profile often attached to this experiment belongs to a different one, Steve the librarian.) The sketch feels informative; the ratio feels like a technicality attached to the question rather than a fact about the answer.

The mechanism is representativeness: a description resembling a stereotype is judged probable in proportion to how well it resembles it, not to how often that category occurs. Frequency data feels like it belongs to a different question than the one in front of you; a specific case feels self-contained. This is a claim about reasoning, not about instruments — for how a test's own error rates interact with a population's base rate, see Sensitivity and Specificity.

The fallacy recurs anywhere a rare-category judgment is made from a compelling profile: forecasting, screening, and Automation Bias toward a system's specific output over its known error rate. The Monty Hall Problem is a neighbour rather than an instance: it turns not on a neglected base rate but on failing to condition on the host's constrained choice of door, which carries information. Stating the base rate explicitly, as a habit rather than an afterthought, is most of the fix.

See also5

Hand-picked in the note itself — the neighbours worth reading next.

Related3

Nearby in the graph rather than deliberately chosen. Looser, sometimes surprising.

Linked from2

Notes elsewhere in the wiki that reach for this one.