3 min read

48 Groups, 104 Games, One Spreadsheet: Grading a World Cup Prediction Model

footballWorld Cupanalyticsmodelling

Thirty-five out of forty-eight.

That's how many group-stage finishing positions the model called correctly at the 2026 FIFA World Cup, before a ball was kicked. Eleven of twelve group winners. All four semi-finalists identified in advance. And a final match-pick record of 51 wins, 48 losses, and 5 pushes across the 104-game tournament: 52 per cent of decided games, which is roughly what you'd expect from a model honest enough to admit that football is harder to predict than it looks.

The pipeline was built for SavvyPlays and ran autonomously on a server in Germany throughout the tournament. Every match got a preview (48 group-stage games and every knockout fixture) generated from a combination of ELO ratings, squad strength metrics, and historical tournament performance. Odds were captured from four bookmakers every thirty minutes. Results and match events were ingested automatically within minutes of the final whistle.

The discipline was in what the model didn't do. It never adjusted its ratings mid-tournament to chase results. It never overrode a pick based on narrative. France, the pre-tournament favourite and the model's champion pick, lost to Spain in the semi-final. The model took the loss and graded it. Every prediction was timestamped and published before kickoff, so the scorecard is a real record, not a reconstructed one.

Where it worked: the group stage. Seventy-three per cent of group positions correct, 78 per cent of advance/eliminate calls right. The model understood tournament football's structure. Weaker teams in groups of death get eliminated regardless of individual match margins, and the model's ELO tiers captured that sorting function well.

Where it didn't: match-level picks in the knockouts, and the Asian Handicap market specifically (1 win, 6 losses, 4 pushes). The AH result was the tournament's clearest lesson: a model built on match-winner probabilities has no business expressing opinions about winning margins in single-elimination games where 120-minute grinds and penalty shootouts dominate. That market was retired from the picks at the review stage.

The champion pick, France, was wrong. Spain won, beating Argentina in extra time through a Ferran Torres goal. The model had France and Spain in the same half and picked France through; Spain's route was the one it gave lower probability to. The four semi-finalists were correctly identified, but the ordering within them wasn't. At a tournament level, calling the final four from a field of 48 is a genuine result. Missing the winner from those four is a genuine miss.

The golden boot shortlist included Mbappé, who finished with a record ten goals. The model's odds-implied favourite list had him joint-top entering the knockouts.

A tournament model's job isn't to be right on every game. No model can be, and anyone claiming otherwise is selling something. Its job is to be calibrated: right about as often as it says it will be, up front about what it doesn't know, and transparent enough that the scorecard is published alongside the predictions. This one was.

The full tournament review, scorecard, and methodology are at savvyplays.com/world-cup.