The wheel of fortune produced a bias; the same experiment produced a heuristic
Spin a rigged wheel in front of a research participant and let it stop, apparently at random, on the number ten or the number sixty-five. Ask the participant, right after, what percentage of African countries belong to the United Nations. The number the wheel produced has no bearing on the answer, and everybody who takes the test knows it. People asked after a low spin give lower estimates than people asked after a high spin anyway. Amos Tversky and Daniel Kahneman reported this result, among many others, in their 1974 Science paper “Judgment under Uncertainty: Heuristics and Biases,” and used it to name a heuristic they called anchoring and adjustment: judgments under uncertainty are made by starting from a reference point and adjusting away from it, and the adjustment is characteristically insufficient, so the final answer stays too close to the starting value even when the starting value is a number from a spinning wheel [1]. The paper is the founding document of a research program, and it is worth reading for what it actually says rather than for the slogan it became. Tversky and Kahneman described three heuristics — representativeness, availability, and anchoring — each defined operationally, each demonstrated with a specific quantitative departure from a specific normative model: the engineer-lawyer problem, in which subjects given an uninformative personality sketch ignored a stated base rate of seventy versus thirty percent engineers and judged the probability at fifty-fifty regardless of which base rate they had been told, a violation of Bayes’ rule the authors could compute exactly [1]. The hospital problem, in which most subjects judged a small maternity hospital and a large one equally likely to record a day with more than sixty percent male births, when elementary sampling theory says the small hospital should stray from fifty percent far more often [1]. The urn problem, in which subjects treated a sample of five balls with four red as roughly as informative as a sample of twenty balls with twelve red, though the correct posterior odds — eight to one against sixteen to one — differ by a factor of two [1]. Every one of these is a claim with a named normative model attached, and the departure from that model is the entire empirical content of the heuristics-and-biases program: intuitive judgment is scored against the calculus a fully informed, unlimited-time reasoner would use, and the size and direction of the gap is the finding.
Gerd Gigerenzer’s program does not deny any of these results. It denies that the calculus used to score them is the only one a mind could reasonably be held to. Gigerenzer and Daniel Goldstein’s 1996 Psychological Review paper “Reasoning the Fast and Frugal Way” opens by naming the target directly: much of contemporary research, they write, treats probability theory and logic as “both the normative and the descriptive models for inference and decision making,” so that a discrepancy between the two automatically counts as an error, and the heuristics-and-biases program “concluded that human inference is systematically biased and error prone” on exactly that basis [2]. Their alternative, following Herbert Simon’s bounded rationality, is to ask what a mind that must decide with limited time, limited knowledge, and limited computational capacity should do, and to test whether the resulting “fast and frugal” algorithms — ones that neither search for all available information nor integrate all of it once found — can match the accuracy of the textbook alternative on real-world tasks rather than on textbook urns [2]. This is a genuinely different research question, not a restatement of the first one with kinder adjectives, and the difference is what the rest of this article follows.
The two programs collided directly and publicly in 1996, in the same issue of Psychological Review, and the collision is worth reporting as an argument rather than as gossip. Kahneman and Tversky’s “On the Reality of Cognitive Illusions” opens by naming Gigerenzer’s charge precisely: that “biases are not biases” and that the heuristics invoked to explain them are “meant to explain what does not exist,” because a normative foundation built on subjective probability is, on Gigerenzer’s frequentist objection, not a foundation at all [4]. Kahneman and Tversky’s response was not a philosophical rebuttal of frequentism; it was empirical. Their central counter-claim is that judgments of frequency — the very format Gigerenzer’s program treats as the mind’s native currency — are “susceptible to large and systematic biases” in their own experiments, so that recasting a problem in frequency terms does not, by itself, make the errors disappear [4]. They also pushed back hard on the scope of Gigerenzer’s critique, pointing out that only two of the twelve biases in their original 1974 paper concern subjective probability at all; the other ten, including anchoring itself, illusory correlation, and misperception of randomness, are not touched by an objection to single-event probability, and Gigerenzer’s account of the field, they wrote, does not mention them [4]. Gigerenzer’s same-issue reply, “On Narrow Norms and Vague Heuristics,” held that Kahneman and Tversky had overstated his claim into “all cognitive illusions disappear in frequency formats,” when his actual and more limited position was that he had proposed and tested specific models predicting when frequency judgments are valid and when they are not [5]. Both sides, in other words, accused the other of arguing against a simplified version of their position — a normal and often justified complaint in an academic dispute, and one that leaves the substantive question exactly where it started: under which descriptions of a problem does human judgment approximate a normative model, and under which does it not.
Natural frequencies did to base-rate neglect what a units conversion does to long division
The heuristics-and-biases program’s strongest single result concerns a physician’s task: a patient tests positive on a screening exam, and the physician must estimate the chance the patient actually has the disease. Posed as probabilities, the mammography problem reads: the probability a woman age forty has breast cancer is one percent; if she has cancer, the probability she tests positive is eighty percent; if she does not, the probability she tests positive anyway is nine and six-tenths percent. What is the probability that a woman with a positive test actually has cancer? Applying Bayes’ rule gives seven and eight-tenths percent. Ward Edwards found decades earlier that intuitive answers to problems of this shape are “conservative” relative to the Bayesian value, and researchers going back to the 1970s concluded that people neglect the base rate almost as a matter of course; Robyn Dawes’s phrase for the state of the field was that base-rate neglect’s “gruesomeness, robustness, and generality… are matters of established fact” [3]. David Eddy’s 1982 survey of one hundred physicians found ninety-five of them estimated the posterior probability at between seventy and eighty percent — roughly ten times the correct value — when given exactly this problem in probability form [3].
Gerd Gigerenzer and Ulrich Hoffrage’s 1995 Psychological Review paper argues that the deficit is not in Bayesian inference as such but in the information format the standard experiments use. Recast the same facts as natural frequencies — the format a forager or a physician without a pocket calculator would actually encounter by observing cases one at a time — and the arithmetic changes shape: ten out of a thousand women this age have breast cancer; of those ten, eight will test positive; of the nine hundred and ninety without cancer, ninety-five will also test positive. A new woman tests positive — how many of the women like her, in this frequency, actually have cancer? The answer is eight out of a hundred and three, which is the same seven-point-eight percent, but the calculation needed to reach it no longer requires normalizing three conditional probabilities against a base rate the way Bayes’ formula does in its textbook form; it requires only comparing two counts that are already directly comparable, because natural sampling never separates the numbers from the population they came out of in the first place [3]. Gigerenzer and Hoffrage’s theoretical claim, stated as a formal result rather than a rhetorical one, is that the same posterior probability can be computationally cheap or computationally expensive to derive depending entirely on the format the input arrives in — mathematically equivalent, not psychologically equivalent [3]. Across the studies they ran, naive undergraduate participants derived the Bayesian answer via some identifiable algorithm — full calculation, a pictorial “beam” analog they invented on the spot, or one of several simplifying shortcuts the authors catalogued — in up to half of all the problems they were given, when those problems were phrased as frequencies, a proportion far above what probability-format studies had found up to that point [3].
The improvement is not confined to undergraduates solving puzzles. Gigerenzer reports asking a continuing-education session of one hundred and sixty practicing gynaecologists the mammography question directly: a majority, sixty percent, believed the answer was between eighty and ninety percent, and nineteen percent believed it was one percent — a spread that, on his account, would rightly frighten a patient relying on her doctor’s arithmetic [18]. After a single session teaching the natural-frequency translation, most of the same gynaecologists — eighty-seven percent — had mastered converting sensitivities and false-positive rates into frequencies and computing the positive predictive value correctly [18]. A separate study of forty-eight physicians in Munich and Düsseldorf, asked to estimate the positive predictive values of four different diagnostic tests, found accurate estimates were rare when the tests were described in conditional-probability form and reliably more common once the same information was translated into natural frequencies [19]. None of this shows that base-rate neglect is a myth; the Eddy result and the sixty-percent-plus-nineteen-percent gynaecologist result are themselves demonstrations that the probability format defeats most people most of the time. What the frequency-format results show is that the deficit is specific to a representation, not a fixed property of the reasoning mechanism itself, which is exactly the distinction Kahneman and Tversky’s 1996 reply denied applied generally.
Not knowing a city can be worth more than knowing it, but only sometimes
The recognition heuristic is the starkest demonstration in the ecological-rationality program, because it produces a result that looks, on first encounter, like an error: less information yielding a more accurate inference than more information. Daniel Goldstein and Gigerenzer’s 2002 Psychological Review paper “Models of Ecological Rationality” states the heuristic plainly — if one of two objects is recognized and the other is not, infer that the recognized object scores higher on the criterion — and reports the demonstration that made the case. Asked which city has the larger population, San Diego or San Antonio, a group of American undergraduates who knew both cities answered correctly about two-thirds of the time. A comparable group of German undergraduates, many of whom had never heard of San Antonio, answered correctly one hundred percent of the time, because for them the question reduced entirely to whether they recognized a name, and recognition happened to correlate almost perfectly with population size in that pair [6]. Goldstein and Gigerenzer report a second version of the same asymmetry with English football: fifty Turkish students and fifty-four British students forecast the outcomes of thirty-two English FA Cup third-round matches. The Turkish students, who recognized some of the English teams and not others, forecast sixty-three percent correctly; the British students, who knew the league intimately, forecast sixty-six percent correctly — nearly the same accuracy from a group with dramatically less knowledge, because the Turkish students’ guesses tracked which team name they recognized in six hundred and twenty-seven of six hundred and sixty-two informative comparisons, or ninety-five percent of the time [6]. The formal claim behind the demonstration is that a recognized object beats an unrecognized one whenever the recognition validity — the empirical correlation between having heard of something and its actually scoring higher — is strong enough, and that adding more knowledge about the recognized items can, past a point, decrease accuracy rather than raise it, an effect Goldstein and Gigerenzer name the less-is-more effect [6].
The most consequential test of the recognition heuristic put money behind it. Bernhard Borges, Goldstein, Andreas Ortmann and Gigerenzer’s chapter “Can Ignorance Beat the Stock Market?” built portfolios out of company names that laypeople recognized and reported that these recognition-based portfolios matched or outperformed major mutual funds, market indices, randomly selected stocks, and portfolios of low-recognition stocks, in both German and American markets over the study’s window [7]. The claim travelled further than the original evidence base, and later tests have not been kind to it as a general strategy. Patric Andersson and Tim Rakow’s 2007 replication, “Now You See It, Now You Don’t,” tested the recognition heuristic as a stock-picking rule across multiple markets and periods and found “no support for the claim that a simple strategy of name recognition can be used as a general strategy to select stocks that yield better-than-average returns”; recognized-name portfolios performed no better than unrecognized ones overall, though the authors did find a conditional pattern in which recognition portfolios did comparatively better in falling markets and worse in rising ones [8]. Rüdiger Pohl’s 2006 paper “Empirical Tests of the Recognition Heuristic” raises a related boundary-condition problem even outside finance: across four paired-comparison experiments, people’s individual rate of choosing in line with the recognition heuristic did not track their individual recognition validity for that domain, and additional knowledge about a recognized object changed how often people relied on recognition alone rather than leaving the heuristic’s use untouched [9]. The honest summary, consistent with the house rule against picking a winner the evidence has not picked, is that the recognition heuristic is a real and well-specified phenomenon in the domains where it was first demonstrated — city populations, sports outcomes — and an unreliable general-purpose trading strategy once tested against harder, adversarial markets where recognition itself is partly a product of the very prices being forecast.
One good reason beat multiple regression in a fair fight
Ecological rationality’s least intuitive claim is not about recognition specifically but about integration: that a decision rule using only the single most valid cue it can find, and stopping the moment that cue discriminates between two options, can match or beat a rule that weighs and combines every available cue. Gigerenzer and Goldstein’s 1996 paper tested this directly with an algorithm they called Take The Best, applied to the standard drosophila task of their research program: comparing the populations of German cities using nine binary cues such as whether each city has a professional football team or a university. Take The Best searches cues in order of validity and stops at the first one that distinguishes the two cities, ignoring everything after it; the “rational” competitors in their simulation — multiple regression given complete information about all eighty-three cities, plus tallying, weighted tallying, and two linear models — used every cue, weighted and combined [2]. Averaged across every level of simulated knowledge, from ten percent of cue values known up to complete knowledge, Take The Best drew correct inferences sixty-five point eight percent of the time, matching weighted tallying’s sixty-five point eight percent and edging out multiple regression’s sixty-five point seven percent, while looking up on average fewer than six of the twenty possible cue values it could have checked [2]. The two linear models that combined recognition with every other cue scored markedly worse, around sixty-two percent, because they let strong negative evidence on other cues override a recognized-versus-unrecognized signal that was, in this environment, close to eighty percent valid on its own [2]. Correct inferences peaked, across every algorithm, at around seventy-seven percent, and — again demonstrating the less-is-more effect independently of recognition — that peak was not reached at full knowledge of all cue values but at an intermediate level, after which learning more about already-recognized cities could make an already-limited-knowledge participant’s inferences worse rather than better [2]. Gigerenzer, Goldstein and their colleague Jean Czerlinski later extended the same competition across twenty real-world prediction environments unrelated to city populations, training each strategy on half of each dataset and testing it on the other half; Take The Best’s advantage held up under this harder, out-of-sample test against multiple regression, which is the comparison that actually matters for judging whether one-reason decision making generalizes rather than merely fitting one convenient environment.
Selection does not care whether a threshold is unbiased, only whether it is cheap
The heuristics-and-biases program treats a systematic departure from an unbiased estimator as the phenomenon to be explained. Error management theory, developed by Martie Haselton and David Buss, supplies a reason a biased estimator can be exactly what natural selection should be expected to build. The argument is a direct application of signal detection theory to fitness rather than to accuracy. Consider a judgment made under genuine uncertainty, with two possible errors: inferring a fitness-relevant state is present when it is not — a false alarm — and failing to infer it when it actually is — a miss. Haselton and Andrew Galperin’s later review states the theory’s scope conditions plainly: a bias is expected to evolve wherever a decision recurrently affected fitness, was made under uncertain information, and carried false-positive and false-negative errors whose costs were recurrently asymmetric over evolutionary time [11]. Detecting a snake in the grass satisfies all three: a missed snake can kill, while a false alarm over a stick costs only a moment’s added caution, so a threshold tuned to treat ambiguous shapes as more snake-like than the raw evidence warrants pays for itself in expected fitness even though it generates more false alarms than an unbiased detector would [11].
This is a rule about optimal thresholds, not merely an intuition, and it has a compact formal statement. A decision-maker observes ambiguous evidence
When the two error costs are equal, the cost ratio on the right collapses to one, and the rule reduces to the ordinary Bayesian classifier that a norm blind to consequences would prescribe — the classical, “unbiased” threshold heuristics-and-biases research uses as its comparison point. Error management theory’s substantive claim is that this special case almost never held over evolutionary history for fitness-relevant judgments: whenever
A forager’s caution about a certain meal is a human’s caution about a certain gain
The asymmetric-cost logic behind error management theory has an older and independently derived cousin in behavioral ecology: risk-sensitive foraging theory, which predicts that an animal’s preference for a fixed versus a variable food reward should flip depending on whether the fixed option meets its energy requirement. Thomas Caraco, Steven Martindale and Thomas Whittam’s 1980 experiment with yellow-eyed juncos tested this directly. In the first condition, birds chose between a perch that reliably delivered enough seed to meet the day’s full energy budget and a perch that delivered a variable amount averaging the same; all five test birds preferred the reliable perch, avoiding the gamble even though it offered no worse an average outcome [13]. In the second condition, the researchers shrank the reliable option below what the birds needed to survive the day while leaving the variable option’s chance of abundance intact; four of the five birds reversed their preference and switched to the gamble, including individuals who had chosen the safe option in the first condition, exactly tracking the model’s prediction that risk preference depends on where the guaranteed option sits relative to a survival-relevant reference point rather than on a fixed, general appetite for risk [13].
The shape of that result — risk aversion above a reference point, risk seeking below it — is also the shape of prospect theory’s reflection effect and, by extension, of loss aversion, the observation that losses relative to a reference point are typically weighted more heavily than equivalent gains. Two research traditions arrived at the same S-shaped preference pattern from opposite directions: one from a psychophysics of human choice under uncertainty, the other from an energy-budget model of what a starving forager should rationally do. The convergence is a genuine point in ecological rationality’s favor: a reference-dependent value function is not merely a description of human quirks but a solution an evolutionary process independently found for animals with no capacity for a utility calculation at all.
The twist is that the very asymmetry risk-sensitive foraging theory is invoked to rationalize may be smaller and less general in humans than the textbook account of loss aversion suggests. David Gal and Derek Rucker’s 2018 Journal of Consumer Psychology paper “The Loss of Loss Aversion” reviews the accumulated experimental record and argues that current evidence does not support the claim that losses, on balance, loom systematically larger than equivalent gains across the range of contexts in which loss aversion is routinely invoked; where a loss-framed effect does appear, Gal and Rucker contend it is often better explained by other factors specific to the study design than by a general asymmetry in the value function [14]. The paper’s own publication as a Journal of Consumer Psychology research dialogue signals that this is a live and contested reading rather than a settled overturning: other researchers in the same exchange maintain that a more contingent, context-dependent version of loss aversion survives Gal and Rucker’s critique. Either way, the exchange means an adaptationist explanation should not be built on a bias whose size and universality is itself under active dispute — a caution the ecological-rationality program, with its own emphasis on precisely specifying when an effect holds, should be the first to take seriously.
Anchoring and framing held up under replication; a celebrated pattern-perception finding did not
Neither research program gets to claim the reproducibility high ground, and stating the record for both is more useful than a scorecard that awards one side a clean win. Many Labs 2, Richard Klein and more than one hundred and eighty co-authors’ 2018 preregistered replication of twenty-eight classic and contemporary findings across roughly one hundred and twenty-five samples and over fifteen thousand participants in thirty-six countries, found the classic Asian-disease framing effect from Tversky and Kahneman’s own research program replicated robustly and consistently across sites, with a pooled effect size around g equals zero point four four [16]. A different anchoring manipulation tested in the same battery — an incidental, environmentally present numeric anchor’s effect on an unrelated consumer estimate, distinct from the classic arbitrary-number-then-estimate paradigm — did not replicate; the pooled estimate across sites was small and statistically indistinguishable from zero [16]. The lesson is not that “anchoring” as a family failed and “framing” succeeded; it is that these are not homogeneous categories, and a result’s replication status has to be reported for the specific paradigm tested, not for the textbook heading it sits under.
A sharper reversal concerns a finding closely related to the heuristics-and-biases program’s account of misperceived randomness — the belief, which Tversky and Kahneman’s original paper already discussed under “misconceptions of chance,” that streaks in a genuinely random process are less likely than they actually are, the same intuition underlying the gambler’s fallacy [1]. Its mirror image, the “hot hand fallacy” in basketball shooting, was declared exactly that — a fallacy, a belief in nonexistent streakiness — in a canonical study from the 1980s. Joshua Miller and Adam Sanjurjo’s 2018 Econometrica paper “Surprised by the Hot Hand Fallacy?” identified a subtle selection bias built into the standard statistical test those studies used: estimating the conditional probability that a hit follows a streak of hits, from a single finite sequence, is itself biased toward finding no streakiness even when real streakiness is present, and the size of the bias is large enough, at realistic sequence lengths, to fully explain the original null results [17]. Correcting for it, Miller and Sanjurjo report, reverses the conclusions of the canonical hot-hand studies and their replications [17]. This is a case in which a heuristics-and-biases-adjacent finding was overturned not by a failed replication in the ordinary sense but by a statistical correction to the original analysis — arguably a more serious kind of correction, since every dataset built on the same flawed estimator inherited the same error. Ecological rationality’s side of the ledger has its own unresolved cases rather than a clean record: the recognition heuristic’s stock-market application did not survive Andersson and Rakow’s test, and Pohl’s boundary-condition work complicates even the city-population demonstrations that motivated the theory in the first place [8, 9]. Robust results exist on both sides, contested and reversed results exist on both sides, and a reader who trusts either program’s replication record more than the other’s has not looked closely enough at either one.
A model of bounded optimality does not choose a side, it computes one
The clearest formal reconciliation of the two programs treats both as partial descriptions of the same underlying problem: how should a system with finite computational resources allocate them. Falk Lieder and Thomas Griffiths’s 2020 Behavioral and Brain Sciences target article “Resource-Rational Analysis” proposes exactly this unification, arguing that human cognitive mechanisms — including the heuristics both research traditions study — are best understood as the optimal use of limited time, memory, and computation rather than as either flawless probabilistic inference or as an arbitrary bag of environment-specific tricks [15]. On this account, a heuristics-and-biases result is real: relative to an idealized reasoner with unlimited computation, human judgment does deviate, sometimes badly, and the deviation is exactly what the classic experiments measure. An ecological-rationality result is also real: many of those deviations are not random noise but the predictable output of a system solving a harder, more honest problem — get a good enough answer, using the cues actually available, within a real budget of time and mental effort — and once that budget is priced into the objective function, the “biased” heuristic can be shown to be the rational solution to the resource-constrained problem rather than an approximation failure of the unconstrained one.
Resource-rational analysis does not settle every dispute the two programs have had; it reframes what counts as evidence for each side. A finding that a heuristic departs from the unconstrained normative model is no longer, by itself, evidence that the mind is malfunctioning, because the unconstrained model was never the one the mind had to solve. But a finding that a heuristic is ecologically well matched to some environment is likewise not, by itself, evidence that the environment in question is the one that mattered across evolutionary history, or the one a modern experimental task actually recreates — the recognition heuristic’s failure as a general stock-picking rule is a demonstration of exactly this second kind of overreach. The honest version of the reconciliation gives both traditions a job: heuristics-and-biases work specifies what a boundedly rational system had to trade off against, and ecological-rationality work specifies which environments the trade-off was actually calibrated against.
Irrational is a claim about the wrong environment as often as it is a claim about the mind
The most durable lesson from thirty years of argument between these two programs is not that one of them won. It is that “irrational” turned out to be a claim with two different objects, and researchers on both sides sometimes forgot which one they were making. A judgment can depart from an unconstrained normative model and still be exactly what a resource-limited system calibrated to a specific environment should produce; whether that departure counts as a bug or a feature depends entirely on whether the environment presenting the judgment is the one the mechanism was actually tuned to. The recognition heuristic is not “rational” or “irrational” as a general fact about minds; it is well calibrated to city-population and sports-outcome judgments, where name recognition tracks the criterion strongly, and poorly calibrated to adversarial financial markets, where recognition is itself partly a product of the price movements a trader is trying to forecast. Error management theory’s asymmetric thresholds are not evidence of a broken detector; they are evidence of a detector priced correctly against the two error costs it evolved to face, which is a different claim from the claim that the detector is accurate.
That reframing yields a specific, checkable prediction for where the field goes over roughly the next decade rather than a truce. On the assumption that resource-rational analysis and its successors continue to gain traction as a unifying framework, expect a shift in how positive results in judgment-and-decision-making journals are framed: fewer papers claiming a heuristic is simply “biased” or simply “adaptive” in general, and more papers specifying the environment — the cue structure, the error-cost asymmetry, the time budget — under which a given heuristic is or is not well calibrated, with the environment treated as an explicit, testable part of the claim rather than left implicit. The observable indicator is directly countable: the proportion of judgment-and-decision-making publications that state an explicit environmental scope condition for a heuristic’s performance, rather than reporting accuracy or bias as a context-free property of the heuristic itself. The disconfirmation condition is equally direct — if a decade from now the field’s flagship journals are still running the debate in the same context-free terms, asking whether human judgment “is” biased or “is” adaptive rather than asking under which environments each mechanism performs well, then the reconciliation this article has described will have been a paper exercise that working researchers ignored rather than a genuine change in how the question gets asked.