AIAI vs the World Cup

Analysis · GPT vs Gemini

GPT vs Gemini: near-identical brackets, decided on the margins

Based on the committed predictions, locked under hash b233af4f.

Put GPT's and Gemini's headline picks side by side and they are almost a carbon copy. Same champion in France, same beaten finalist in Spain, same Golden Boot in Mbappe, and the same playmaker in Michael Olise. So how does a leaderboard ever separate them?

Where they line up

On every marquee call, GPT and Gemini agree. Both make France the champion in their awards form and carry France through the bracket as well. Both put Spain in the final. Both back Mbappe for the Golden Boot and Olise as the leading creator. If you only read the top line, these two models look interchangeable.

Why that does not make them the same

The headline picks are the smallest part of a model's output. Behind each match sits a full probability distribution, a home win, draw, and away win number that has to add up to one. Two models can both pick the same winner while disagreeing sharply on how likely that win is, and on how the remaining probability is split between a draw and an upset.

That is where GPT and Gemini will come apart. The leaderboard does not score who you picked, it scores how well your stated confidence matched what happened. A model that says a France win is 70 percent likely and a model that says it is 55 percent likely have made the same pick and very different forecasts. Over 72 group games, those gaps accumulate into a real difference in Brier score even if the two models never once disagreed on a winner.

The quiet contest

This is the most underrated rivalry in the field precisely because it is invisible at the top line. There is no dramatic fork like Claude's Argentina bracket or Grok's Messi bet. Instead there is a slow, game by game test of calibration between two models that reached the same conclusions by slightly different routes. Whoever was a touch more honest about uncertainty wins, and you will only see it in the numbers.

What to watch

Ignore the picks and watch the average Brier gap between the two on the leaderboard. Since they agree on so much, any separation that opens up is almost pure calibration, the cleanest read you will get on which of the two understood uncertainty better.

Read nextWhat Brier score measuresWhy confidence, not just picks, decides the board. Read nextDeepSeek and calibrationA third model on the same France line. Go toLive leaderboardWatch the Brier gap open up.

Back to the live site