Probability quality
How football model calibration works
A calibrated football model assigns probabilities that match long-run outcome frequencies. If events forecast at about 70% occur about seven times in ten across a suitable sample, that group is well calibrated.
· 8 min read
Calibration connects stated confidence with observed frequency
Calibration asks whether a set of probability forecasts behaves as advertised. It does not ask whether every 70% selection wins. It asks whether comparable events assigned probabilities near 70% occur at roughly that rate over time.
Reliability buckets make calibration visible
A reliability diagram groups forecasts into probability ranges, then compares each group's average forecast with its observed outcome rate. Perfect agreement lies on the diagonal where forecast probability equals observed frequency.
A negative gap can indicate overconfidence in that bucket; a positive gap can indicate underconfidence. The sign and size should be read with the number of forecasts in the bucket.
Worked example: reading one probability bucket
That difference is descriptive evidence, not proof of a stable defect. Sampling uncertainty, competition mix and changes to the model all affect how confidently the gap can be interpreted.
Sample size and composition determine what the chart can support
A bucket with only a few forecasts can move sharply after one result. Wider buckets contain more observations but hide detail; narrower buckets show more detail but become noisier. Counts should therefore appear beside every calibration estimate.
- Use predictions recorded before kickoff.
- Keep settlement rules and markets consistent.
- Show the number of forecasts in each bucket.
- Separate materially different model versions or evaluation periods.
- Report uncertainty instead of treating small gaps as conclusive.
Calibration is not the same as resolution
A model can be calibrated yet uninformative if it gives nearly every event the same base-rate probability. Resolution describes whether forecasts meaningfully separate lower-probability cases from higher-probability cases. Proper scoring rules such as the Brier score assess probability quality while rewarding useful separation.
Calibration must be monitored after deployment
Team strength, competitions, data feeds and model code change. A result from one historical test does not guarantee future calibration. Evaluation should use out-of-sample forecasts and preserve model-version boundaries so deterioration is not hidden by aggregation.
- Compare forecast periods rather than only lifetime totals.
- Investigate persistent gaps, not isolated match results.
- Record corrections without rewriting the original forecast.
- Retest after material input, feature or model changes.
What calibration cannot promise
Good calibration describes a population of forecasts. It cannot identify which individual match will be the exception, remove football variance or guarantee profit at available odds.
Sources and further reading
- A new vector partition of the probability score, Journal of Applied Meteorology
- Strictly proper scoring rules, prediction, and estimation, Journal of the American Statistical Association
- Sampling uncertainty and confidence intervals for the Brier score, Weather and Forecasting