Calibration and the Brier score
Two practical tools for checking whether published probabilities deserve trust—without turning one metric into a trophy.
Calibration asks an intuitive question
Group similar forecasts into probability bands and compare their predicted rate with the observed rate. If 100 events forecast near 60% contain roughly 60 successes, that band is consistent with calibration—subject to sampling uncertainty.
Small samples can look dramatically over- or under-confident by chance. Confidence intervals and cohort size belong beside the chart.
The Brier score measures squared probability error
For a binary event, subtract the outcome—one for occurred, zero for did not occur—from the predicted probability, square the difference, then average across observations. Lower is better.
The score rewards honest probabilities and penalizes confident mistakes, but its meaning depends on event frequency and the benchmark used for comparison.
No single validation metric is enough
FORM/PRICE pairs calibration and Brier score with log loss, market comparison, closing-price coverage and complete prospective records.
- Compare with a simple baseline.
- Keep train and evaluation periods separate.
- Publish missing-price coverage.
- Retain failed and withdrawn decisions.
Sources
Direct links are preserved so the editorial reasoning can be checked independently.