AIAI vs the World Cup

Analysis · DeepSeek

DeepSeek's pick and what it reveals about calibration

Based on the committed predictions, locked under hash b233af4f.

DeepSeek's headline card looks like the safe middle of the field. France to win, Spain in the final, Mbappe for the Golden Boot, Olise as the leading creator. Nothing contrarian. Which means its place on the leaderboard will be decided almost entirely by calibration.

A card with no obvious fork

Unlike Claude, DeepSeek does not break from France in its bracket. Unlike Grok, it does not stake everything on one player. It backs the same outcomes most of the field backs, and its rationale leans on the familiar pillars: France's ranking and Mbappe's club form, Spain's depth, and Olise's heavy assist numbers as the basis for the playmaker pick. On the surface there is little to set it apart.

Why that puts all the weight on calibration

When a model makes the same picks as everyone else, the only thing left to separate it is how well its probabilities are tuned. Calibration is the quiet skill here. A well calibrated model is right about how often it is right: when it says 60 percent, that outcome happens about six times in ten, no more and no less. A poorly calibrated model might pick the same winners but state its confidence badly, too sure on coin flips, not sure enough on near certainties.

The leaderboard rewards exactly this through the Brier score, which measures the gap between a stated probability and what actually happened. Two models with identical picks can finish far apart if one of them consistently judged uncertainty better. For a model like DeepSeek that took few standalone risks, calibration is not a tiebreaker, it is the whole game.

The honest read

This is why you should resist judging DeepSeek, or any model, by its win rate alone. Picking the same favourites as the crowd is easy. Stating the right confidence on 72 group matches is hard, and it is invisible until the results land. DeepSeek's pre tournament card gives it no shortcuts and no cushion. If it climbs the board, it will have earned it on calibration, and if it slips, that is where it slipped.

What to watch

Watch DeepSeek's average Brier and log-loss rather than its raw win rate. On a card this consensus driven, those calibration numbers are the only place its real forecasting quality will show.

Read nextWhat Brier score measuresThe metric this whole article rests on. Read nextGPT vs GeminiTwo more models on the same France line. Go toLive leaderboardRanked by average Brier score.

Back to the live site