Designing Decisions / Essay 06

Statistical Memory in Decision Making

Plans fail because decision makers believe in uniqueness. Reference classes reveal sameness.

12 April 20256 min readDecision Science and Behavioural Design

Introduction: The Illusion of Uniqueness

In February 2022, a clear consensus spread across international affairs circles. The Russian invasion of Ukraine would be swift. Many analysts claimed Kyiv would fall in three weeks. Some forecasts even offered three days. The consensus rested on overwhelming asymmetry in troop numbers, air power and economic might. These predictions were delivered with confidence because they were rich in operational detail and detail often feels like understanding. Yet as the conflict stretched into its third year it became clear that the forecasts did not align with reality.

This prediction failure was not an isolated singular political misjudgment. It was a systemic cognitive error rooted in the Inside View, a mode of reasoning that builds predictions from the unique features of the case at hand. Russian planners focused on tank columns and assumed air dominance. Western analysts focused on GDP ratios, early troop movements and weather constraints. Different details, same method. Both sides treated the moment as a one-off that demanded a one-off forecast.

Humans are easily seduced by granularity. High-fidelity information creates an illusion of control over uncertainty. In addition, when a situation feels distinctive, decision makers overweight details and underweight the variable that often carries the strongest signal, the base rate of similar events. That omission invites optimism bias and the uniqueness fallacy. It also weakens institutional memory, because what looks unprecedented does not get checked against a distribution. The inside view replaced the statistical structure that could have grounded expectations.

Two-panel illustration: a figure releasing a spark beside layered statistical ridges.
Two-panel illustration: a figure releasing a spark beside layered statistical ridges.

Detail as a Substitute for Probability

Reference Class Forecasting (RCF) disciplines judgment by refusing to treat the current situation as unique. Reference class forecasting forces the prediction to begin with observed outcomes from comparable cases, then permits local detail to move the estimate at the margin. Its core function is simple. It replaces storytelling with empirical anchoring.

The method has an order of operations. First, blind the details. Set aside the narrative features that make the present case feel special. Second, select the reference class. Identify the historical set of comparable events, such as major power invasions against sovereign states. Third, anchor the belief. Extract the base rate distribution of outcomes in that class and treat it as the prior.

A useful way to picture the outside view is as a gravity field. The distribution of past outcomes pulls predictions back toward what tends to happen, even when the current narrative argues for exception. Detail still matters, but detail becomes a refinement rather than a foundation. The discipline is less about being pessimistic and more about being calibrated.

Formally, reference class forecasting selects a historical set, such as war duration measured in months:

R={X1,X2,...,Xn}R=\left\{ X_{1}, X_{2}, ..., X_{n} \right\}

It then derives the empirical distribution as:

FR(x)=1ni=1n1(xix)F_{R}(x)=\frac{1}{n}\sum_{i=1}^{n} 1({x_i \leq x})

FR(x) tells you what fraction of comparable cases ended by time x. A heavy upper tail is a warning that “quick resolution” is not the default.

Applied to Ukraine, a reference class of asymmetric post-1945 conflicts would likely have revealed a wide spread of durations, including a long upper tail. That tail matters because it represents the scenarios that strain budgets, alliances, supply chains and domestic political patience. History shows that these wars are almost never short. They get bogged down by resistance, logistics and political mobilization. In strategy, tails are often where systems break.

When the Past Is Chosen Poorly

Reference class forecasting is powerful, but it is not automatic. The forecast is only as good as selecting the correct class. Some analysts probably consulted history in early 2022 but anchored on the wrong reference class. A common comparison frame emphasized post-Soviet fragility: corruption, slow modernization and brittle institutions. Within that class, the probability of rapid collapse can look high.

Chart comparing cumulative failure curves for post-Soviet fragility and existential defence over time.
Chart comparing cumulative failure curves for post-Soviet fragility and existential defence over time.

The problem is that Ukraine after 2014 was not identical to Ukraine before 2014, and wartime behavior is often a class shift, not a parameter tweak. After the Euromaidan revolution, the annexation of Crimea and the outbreak of war in the east, Ukraine had introduced reforms that seemed marginal but meaningful in command structure, decentralization and civil defense. More importantly, when the invasion began, the system moved into a different class: nations fighting for existential survival. In that class, resistance and adaptation are no surprise. In that class the historical base rate of resistance is substantially higher:

P(Reform SuccessPost Soviet Class)LowP(\text{Reform Success} \mid \text{Post Soviet Class}) \approx \text{Low}

The hybrid failure emerged because the analysts identified the class for pre-invasion reform but not the class for wartime resilience. Decision makers reach for the right tool but anchor it to an inaccurate comparison set. The math supplies structure but human judgment supplies classification. Decision makers need to notice when systems jump classes because distributions change when incentives and identity change.

Inside View Errors Beyond War

The inside view does not belong to war alone. Consider a corporate product launch. A leadership team can build an air-tight forecast from the particulars, new features, channel strategy and a confident rollout calendar. The plan feels unique because the product feels unique. Yet the base rate for complex launches, especially those that require ecosystem coordination, often includes delays, demand misreads and second-order constraints that were invisible during the planning phase. Reference classes do not remove ambition. They prevent ambition from being mistaken for probability.

Public policy has the same structure. Major infrastructure programs and health system reforms are frequently forecast by narrative, staffing plans, procurement timelines and political assurances. The outside view asks a different question. How often do comparable reforms deliver on schedule given the procurement regime, the institutional capacity and the incentive environment? The answer is usually a distribution, not a date.

This is the institutional point. Inside-view forecasts are rewarded because they produce compelling narratives with decisive timelines. Outside-view forecasts are discounted because they sound like averages. Yet averages and distributions often carry more information than stories built around the present.

Illustration of a silhouette against concentric rings representing an accumulating record.
Illustration of a silhouette against concentric rings representing an accumulating record.

Conclusion: From Storytelling to Empirical Grounding

A more disciplined question in 2022 would not have been whether Kyiv would fall in three weeks. It would have been how often invasions in the relevant reference class succeed in three weeks and what the distribution says. That shift replaces narrative plausibility with empirical accountability.

The deeper failure was institutional as much as cognitive. Many organizations reward narrative mastery rather than empirical grounding. Institutions reward the analyst who can forecast with confidence. That confidence reads as competence, even when it is uncalibrated. Historical averages feel uninspiring by comparison because they lack drama, but they are often the only honest starting point.

The reform is practical. Leaders should require reference class assessments as a standard input to major decisions. Analysts should be evaluated on calibration against distributions, not on the persuasiveness of their narratives. When the class is uncertain, institutions should say so explicitly and test sensitivity across plausible classes. Forecasting is not prediction. It is preparedness under uncertainty. Reference classes do not remove risk. They reduce fragility by forcing judgment to begin where history actually places it.

Infographic titled 'Two Paths to a Prediction', comparing the narrative inside view with the data-driven outside view.
Infographic titled 'Two Paths to a Prediction', comparing the narrative inside view with the data-driven outside view.

Related writing

Shared inquiry fields