FORM/PRICEFootball intelligencePreview the app
← Back to Learn
Model validation · 7 min read

Calibration and the Brier score

Two practical tools for checking whether published probabilities deserve trust—without turning one metric into a trophy.

01

Calibration asks an intuitive question

Group similar forecasts into probability bands and compare their predicted rate with the observed rate. If 100 events forecast near 60% contain roughly 60 successes, that band is consistent with calibration—subject to sampling uncertainty.

Small samples can look dramatically over- or under-confident by chance. Confidence intervals and cohort size belong beside the chart.

02

The Brier score measures squared probability error

For a binary event, subtract the outcome—one for occurred, zero for did not occur—from the predicted probability, square the difference, then average across observations. Lower is better.

The score rewards honest probabilities and penalizes confident mistakes, but its meaning depends on event frequency and the benchmark used for comparison.

03

No single validation metric is enough

FORM/PRICE pairs calibration and Brier score with log loss, market comparison, closing-price coverage and complete prospective records.

  • Compare with a simple baseline.
  • Keep train and evaluation periods separate.
  • Publish missing-price coverage.
  • Retain failed and withdrawn decisions.
S

Sources

Direct links are preserved so the editorial reasoning can be checked independently.

  1. Probabilistic Forecasts, Calibration and SharpnessJournal of the Royal Statistical Society: Series B · research · accessed 31 Aug 2026
  2. The evidence protocolFORM/PRICE · internal · accessed 31 Aug 2026