Skip to content
All writing

Note · Risk & AI

A 70% forecast should be wrong three times in ten

Calibration is the part of forecasting leaders most often misread. What the RGA practicum taught me about judging forecasters on the right thing.

Hizbawi MeresaSeptember 30, 2026 2 min read

Drawn from the case studyRGA Emerging-Risk Monitoring Practicum

What happened

In our RGA practicum, the Metaculus crowd’s calibration error across 6,184 resolved questions was 0.070. In plain terms: when the crowd said 70%, the outcome happened close to 70% of the time.

Why it matters

That also means a well-calibrated 70% forecast is supposed to be wrong about three times in ten. If every 70% call comes true, the forecaster was underconfident — the honest number was higher. Organizations rarely read it that way. They remember the misses, so forecasters drift toward the safe middle, or retreat into words like “likely” and “possible” that can never be scored.

My read

If a team wants better forecasts, it has to change what it rewards. Ask for numbers rather than adjectives. Score them once the outcome is known, in regular post-mortems. And judge forecasters on calibration across many questions, not on the last miss. That was also the logic behind the governance we recommended to RGA: forecasts earn trust by being scored, not by being right once.

Source: RGA × WashU Olin CEL practicum, published results (April 2026).

#Forecasting#Calibration#Decision-Making

The work behind this essay

RGA Emerging-Risk Monitoring Practicum

Assessed whether Reinsurance Group of America (RGA) could rely on Metaculus crowd forecasts for forward-looking risk decisions — assembled a dataset of 11,000+ questions, evaluated forecasts across 6,184 resolved questions in Python, showed 30.5% lower forecast error than baseline, and presented an LLM-augmented workflow to senior stakeholders.

Read the case study