MERIDIAN — public accountability record
Our forecasting record, updated as forecasts resolve.
Every forecast MERIDIAN publishes is scored against what actually happened, and the score is published whether or not it flatters the system.
The headline Brier score with its sample size, calibration broken down by confidence tier, and the full ledger of resolved forecasts, each judged against a resolution criterion set before the fact. The numbers load live from the public API; this page is not a screenshot of a good week.
A Brier score is the mean squared error of probabilistic forecasts: 0.000 is perfect, lower is better. Guessing 50% on everything scores 0.250, so that line is the floor to beat, not a target. The scale is steep below it: a forecaster who is right 80% of the time, and says so, averages about 0.160; right 90% of the time, about 0.090. It rewards being confident and right, and punishes being confident and wrong.
Do our confidence labels tell the truth?
Each forecast is issued in a tier that commits to a probability band — LOW 10–25%, MEDIUM 25–55%, HIGH 55–75%, CRITICAL 75–90%. A calibrated system is one whose observed hit rate lands inside the band it claimed. Comparing predicted against observed per tier says something a single average cannot: where the system is overconfident.
Every forecast is scored on a falsifiable resolution criterion set before the fact. Resolution is performed by an auditor that is blind to the original probability, and the Brier calculation itself is deterministic rather than model-judged. Quantitative thresholds resolve against the underlying data series directly, with no model in the loop at all.
A track record needs volume before it means much. Calibration figures mature around the ninety-day mark; before that, read the sample size next to the score and treat the number as provisional. That caveat is on this page rather than in a footnote because a self-scoring system that hides its sample size is not one.