Essay · Risk & AI
Signal, not oracle: what 6,184 resolved forecasts taught me about trusting the crowd
Our team tested whether a global reinsurer could use crowd forecasts to see emerging risks earlier. The platform beat its baseline by 30.5%. The more useful finding was where that number should — and shouldn’t — be allowed to go.
Hizbawi MeresaSeptember 30, 2026 3 min read Includes an interactive model
In spring 2026, through Olin’s Center for Experiential Learning, our six-person team took on a question from Reinsurance Group of America: could Metaculus, a public crowd-forecasting platform, serve as a forward-looking signal for emerging risks? My role on the team was strategy and forecasting.
The question matters because some of the risks a life reinsurer cares most about — GLP-1 drugs changing longevity, a new pandemic, AI-related liability — can move faster than the experience data actuaries rely on.
The forward-visibility gap
Actuarial science is excellent at pricing what has happened many times before. It is structurally slower with risks that are new. A therapy can begin changing mortality before long-run population data exists. A pathogen can become a portfolio stressor before there is any base rate for that specific scenario.
In that gap, risk often shows up first in public signals — research, regulation, expert judgment — rather than in claims. That is where an external forecasting signal could add value, if it is good enough to trust.
Only score what has resolved
To test that, we assembled a dataset of more than 11,000 Metaculus questions and evaluated the 6,184 binary questions that had already resolved. That restriction matters: you can only score a forecast against an outcome that is known.
The main measure was the Brier score — the squared gap between the probability given and what happened. Across the resolved set, the crowd scored 0.148, a 30.5% improvement over the baseline, with a calibration error of 0.070. Try the scoring yourself:
Score a forecast the way we scored the crowd.
A Brier score is the squared gap between the probability you gave and what happened (1 or 0). Zero is perfect; saying 50% on everything earns 0.25. Lower is better.
Your Brier score
0.090
Good — in the crowd’s league.
Brier score 0.090. Good — in the crowd’s league.
How it compares (lower is better)
- Metaculus crowd · 6,184 resolved questions0.148
- Baseline · historical base rates0.213
- Always answer 50%0.250
- Your single forecast0.090
The crowd’s 0.148 is 30.5% better than the baseline — real skill, measured only on questions that had already resolved.
Calibration check
You gave ten different events a 70% chance. How many actually happened?
Underconfident. Events you called 70% happened 100% of the time — being “right” every time means your 70% should have been higher. The crowd’s calibration error was 0.07: when it said 70%, outcomes landed close to 70%.
What 30.5% does — and doesn’t — mean
A 30.5% skill score means the crowd adds real information beyond historical frequencies. The calibration result means its probabilities are interpretable: when it says 70%, outcomes land close to 70%.
What it does not mean is that a crowd probability belongs in a pricing model. The platform cannot see a reinsurer’s book, underwriting rules, or exposures. It estimates how likely something is, not why it happens or how severe it would be. So our recommendation was deliberately bounded: use crowd forecasts as a signal, not an oracle — a structured trigger to investigate when the outside probability diverges from internal assumptions.
Where LLMs fit
Large language models were not the best forecasters in the research we reviewed. Their value is operational: gathering evidence, retrieving base rates, drafting scenarios, and red-teaming assumptions so an analyst can produce a decision-ready forecast packet faster.
The workflow we proposed kept humans at every decision point: the crowd signal flags a shift, an LLM drafts the evidence summary, an analyst reviews and decides whether to escalate, and anything that could touch pricing goes to actuarial review first. One rule was non-negotiable: no forecast enters pricing, reserving, or capital decisions without that review.
A bounded pilot, not a platform
We did not recommend enterprise integration. We recommended a 90-day pilot with a named owner, a focused set of tracked questions, and explicit go/no-go criteria — judged not only on forecast accuracy but on decision value: did the signal change what the risk team chose to investigate?
What I took from it
Accuracy is necessary, not sufficient. A well-scored signal without governance can create false confidence faster than no signal at all. Governance is part of the product.
Averages hide the useful part. Performance varied by risk domain, so the right playbook differs by domain rather than treating the platform as one number.
Calibration changes how leaders should read forecasts. A well-calibrated 70% should be wrong about three times in ten. Teams that punish every miss push forecasters toward vague language, and vague language can’t be scored.
The work behind this essay
RGA Emerging-Risk Monitoring Practicum
Assessed whether Reinsurance Group of America (RGA) could rely on Metaculus crowd forecasts for forward-looking risk decisions — assembled a dataset of 11,000+ questions, evaluated forecasts across 6,184 resolved questions in Python, showed 30.5% lower forecast error than baseline, and presented an LLM-augmented workflow to senior stakeholders.
Read the case study