Forecast evaluation

How to measure football prediction accuracy

Football prediction accuracy is not one number. Hit rate checks how often a chosen outcome wins; calibration and proper scoring rules check whether the probabilities themselves deserve confidence.

· 9 min read

What does prediction accuracy mean in football?

Prediction accuracy describes how closely forecasts agree with later outcomes, but the right measurement depends on what was forecast. A single match pick can be marked right or wrong. A probability forecast must also be judged by whether its confidence was appropriate.

Accuracy is not one number

Hit rate is the share of picks that win. It is easy to understand, but it discards probability strength and can reward a model for repeatedly choosing the most common outcome. Probability forecasts need measures that keep the full estimate.

  • Hit rate answers how often a chosen class was correct.
  • Calibration asks whether events assigned a given probability occur at about that rate.
  • A proper scoring rule grades every probability and outcome together.
  • Economic results also depend on the prices available, not prediction quality alone.

Worked comparison: the same hit rate, different forecasts

The example shows why a list of winners is incomplete evidence. Model B barely crossed the decision threshold each time, whereas Model A assigned more probability to the outcomes that occurred.

Calibration checks whether confidence matches frequency

Group comparable forecasts into probability ranges, then compare the average forecast with the observed event rate. If selections forecast near 70% occur only 50% of the time across a suitable sample, that group was overconfident.

Calibration gap = observed event rate − average forecast probability

A gap near zero is desirable, but bins can hide variation and small groups are noisy. A calibration plot is more informative when each group’s sample size is shown.

Compare the model with a relevant baseline

A score has little meaning without context. Compare the same matches and market against a simple reference, such as the historical base rate or a consistently margin-adjusted market probability. A model can beat random guessing yet still add no useful information beyond a strong baseline.

  • Use exactly the same fixtures and outcomes for every model.
  • Keep probability definitions and settlement rules identical.
  • Report both the model score and the reference score.
  • Do not compare hit rates from markets with different difficulty or class balance.

Test in time order and outside the training sample

Football data has a time order. Train or tune on earlier matches, then evaluate on later matches whose outcomes were unavailable during development. Randomly leaking future information into training can make a model appear more accurate than it would have been live.

A practical football forecast scorecard

  • Complete count of eligible forecasts and settled outcomes.
  • Hit rate for the declared decision rule.
  • Brier score or another proper score for the probabilities.
  • Calibration by probability range with group sizes.
  • Results against a named baseline on the same sample.
  • Breakdowns by time and market chosen before reviewing results.

Sources and further reading