Probability quality

How football model calibration works

A calibrated football model assigns probabilities that match long-run outcome frequencies. If events forecast at about 70% occur about seven times in ten across a suitable sample, that group is well calibrated.

· 8 min read

Calibration connects stated confidence with observed frequency

Calibration asks whether a set of probability forecasts behaves as advertised. It does not ask whether every 70% selection wins. It asks whether comparable events assigned probabilities near 70% occur at roughly that rate over time.

Reliability buckets make calibration visible

A reliability diagram groups forecasts into probability ranges, then compares each group's average forecast with its observed outcome rate. Perfect agreement lies on the diagonal where forecast probability equals observed frequency.

Calibration gap = observed frequency − average forecast probability

A negative gap can indicate overconfidence in that bucket; a positive gap can indicate underconfidence. The sign and size should be read with the number of forecasts in the bucket.

Worked example: reading one probability bucket

That difference is descriptive evidence, not proof of a stable defect. Sampling uncertainty, competition mix and changes to the model all affect how confidently the gap can be interpreted.

Sample size and composition determine what the chart can support

A bucket with only a few forecasts can move sharply after one result. Wider buckets contain more observations but hide detail; narrower buckets show more detail but become noisier. Counts should therefore appear beside every calibration estimate.

  • Use predictions recorded before kickoff.
  • Keep settlement rules and markets consistent.
  • Show the number of forecasts in each bucket.
  • Separate materially different model versions or evaluation periods.
  • Report uncertainty instead of treating small gaps as conclusive.

Calibration is not the same as resolution

A model can be calibrated yet uninformative if it gives nearly every event the same base-rate probability. Resolution describes whether forecasts meaningfully separate lower-probability cases from higher-probability cases. Proper scoring rules such as the Brier score assess probability quality while rewarding useful separation.

Calibration must be monitored after deployment

Team strength, competitions, data feeds and model code change. A result from one historical test does not guarantee future calibration. Evaluation should use out-of-sample forecasts and preserve model-version boundaries so deterioration is not hidden by aggregation.

  • Compare forecast periods rather than only lifetime totals.
  • Investigate persistent gaps, not isolated match results.
  • Record corrections without rewriting the original forecast.
  • Retest after material input, feature or model changes.

What calibration cannot promise

Good calibration describes a population of forecasts. It cannot identify which individual match will be the exception, remove football variance or guarantee profit at available odds.

Sources and further reading