Probability scoring
Brier score for football predictions
The binary Brier score is the mean squared difference between forecast probabilities and what happened. It rewards probability estimates that are both decisive when justified and close to the observed outcomes.
· 8 min read
What is the Brier score?
For a binary football event such as “home win: yes or no,” the Brier score averages the squared error between each forecast probability and the outcome. Record the outcome as 1 when the event occurs and 0 when it does not. A lower score is better, and 0 is perfect.
The binary Brier score formula
N is the number of forecasts, pᵢ is the forecast probability from 0 to 1, and oᵢ is 1 if the event occurred or 0 if it did not.
Squaring makes every error positive and penalises confident mistakes more heavily. Forecasting 90% for an event that fails contributes 0.81; forecasting 60% for the same failure contributes 0.36.
One forecast: what the score means
A single contribution cannot establish model quality. The score becomes useful when the same rule is applied to a complete series of forecasts.
Worked example across four matches
A constant 50% forecast on those same four balanced outcomes scores 0.25. The model therefore scores better than that particular reference on this small illustrative sample, but four matches are far too few for a stable real-world conclusion.
There is no universal good Brier score
Event frequency and forecast difficulty affect the score. A rare event can give a low score to an uninformative model that nearly always predicts a low probability. Interpret Brier scores against a relevant reference forecast on the identical sample.
Positive values beat the chosen reference, zero matches it and negative values are worse. The conclusion is always relative to that named reference and sample.
The score reflects calibration and resolution
A useful probability system should be calibrated and should separate situations with genuinely different risks. Calibration means that events forecast near a probability occur near that rate. Resolution means the forecasts move away from the base rate when the evidence supports doing so.
Common Brier score mistakes
- Mixing percentage inputs such as 70 with decimal inputs such as 0.70.
- Scoring only published winners and omitting losing or void forecasts.
- Comparing binary and three-outcome definitions as if they were identical.
- Using a reference score from a different set of matches.
- Treating a small improvement as certain without considering sampling uncertainty.
- Reading a good historical score as a guarantee of future profit.
Sources and further reading
- Verification of forecasts expressed in terms of probability, Monthly Weather Review
- Proper scoring rules for estimation and forecast evaluation, Journal of the American Statistical Association
- Sampling uncertainty and confidence intervals for the Brier score, Weather and Forecasting