AIAI vs the World Cup

Analysis · Pre tournament

What five AI models picked before a ball was kicked

Based on the committed predictions, generated 30 May 2026 and locked under hash b233af4f.

On the surface the five models agreed on almost everything. All of them made France the favourite. Dig one layer down, into the brackets and the individual awards, and the picture splits in ways that say a lot about how each model reasons.

The headline: everyone wants France

Ask the five models for a single champion in their awards form and you get the same answer five times: France. That is a rare clean sweep. France went into the tournament as the top ranked side, with recent wins over Brazil and Colombia and with Kylian Mbappé in heavy scoring form, so the agreement is not surprising. What is interesting is how little daylight any of them put between France and the chasing pack. Several rationales describe the final as a coin flip, with phrases like a narrow pick and a low certainty final doing a lot of work.

A clean sweep on the favourite is exactly the kind of pick that does not separate the models. If France wins, everyone banks the same points. The contest will be decided in the places where the models disagreed, which is where this gets fun.

The bracket twist: Claude breaks ranks with itself

Here is the first real split. The awards form is a gut call for the trophy. The bracket is a round by round path that has to survive every matchup along the way. Four models stayed consistent and had France winning both. One did not.

Claude named France in its awards, then played out a bracket in which Argentina lifts the trophy instead. Its reasoning leans on Messi led experience and a favourable draw for Argentina through the knockout rounds. In other words, Claude likes France as a team on paper but, when forced to simulate the actual road to the final, kept landing on Argentina. That tension between a top down favourite and a bottom up path is precisely the sort of thing a bracket exposes and a one line pick hides.

ModelAwards championBracket championAwards runner up
Claude (Anthropic)FranceArgentinaSpain
GPT (OpenAI)FranceFranceSpain
Gemini (Google)FranceFranceSpain
Grok (xAI)FranceFranceArgentina
DeepSeekFranceFranceSpain

Grok is the contrarian on the individual awards

Four of the five models built their picks around the same two players. Mbappé for the Golden Boot, on the back of a 25 goal club season, and a France or Spain creator for the assists crown. Grok went its own way. It named Lionel Messi as both top scorer and top assist maker, and made Argentina its runner up rather than Spain.

The logic is easy to follow. Grok's rationale points straight at Messi's club output, quoting a 35 goal and 23 assist season, and reasons that if Argentina go deep, Messi carries both individual awards. It is a bolder, more concentrated bet. If Argentina flame out early it ages badly, but if they reach the final it could sweep the personal honours while the other four split them.

The quiet disagreement: who makes the assists?

The top scorer pick was near unanimous, with four models on Mbappé and only Grok dissenting. The assist pick was the messiest of all the awards. Three different names came back from five models:

Assists are noisy and hard to forecast, so a three way split is honest. It also gives the leaderboard somewhere to move. Unlike the champion pick, this one will actually separate the models.

What to watch

Because the models clustered so tightly on France, the early scoring will hinge on the group stage probabilities rather than the trophy. That is where small differences in calibration show up first, and where the leaderboard starts to spread. The bracket and the awards only pay off later, but when they do, three picks will matter most:

  1. Does Claude's Argentina bracket look smart, or does it cost the only model that broke from France?
  2. Does Grok's all in Messi bet sweep the individual awards, or collect nothing?
  3. Olise, Yamal or Messi for the assists crown, the one award nobody agreed on.

Every one of these picks is locked and hashed, so none of them can be edited now that the games have started. You can follow how they are holding up on the live leaderboard and dig into individual matches on the predictions page.

Deep dives

Longer reads on each model's pick and the scoring behind the board:

ClaudeThe Argentina bracketWhy Claude broke ranks with its own favourite. GrokAll-in on MessiThe most concentrated bet of the five. GPT vs GeminiNear-identical bracketsSame picks, decided on the margins. DeepSeekA calibration testA safe card whose fate rests on confidence. ConsensusWhere all five agreedAnd why that is the weakest signal. ExplainerWhat Brier score measuresWhy the loudest model is not winning.
Go toLive leaderboardSee who is ahead, ranked by average Brier score. Read nextMethodologyHow the picks were locked and how they are scored. Read nextAboutThe full story behind the experiment.

Back to the live site