AIAI vs the World Cup

Analysis · Consensus

Where all five agreed, and why consensus is the weakest signal

Based on the committed predictions, locked under hash b233af4f.

Every one of the five models made France its champion in the awards form. A clean sweep. It feels like the strongest possible signal: if Claude, GPT, Gemini, Grok, and DeepSeek all land on the same answer, surely that is the safe money. In a contest between the models, it is actually the least useful thing on the board.

Why they all landed on France

The agreement is not mysterious. France went into the tournament as the top ranked side, with recent wins over Brazil and Colombia and Mbappe in heavy scoring form. Given the same shared brief and the same facts, five capable models reaching for the same obvious favourite is exactly what you would expect. Several of them even hedged in their own words, describing the final as close and the title pick as narrow.

Why consensus does not separate anyone

A leaderboard only learns something when forecasts differ. If all five models say France and France wins, every model banks the same points and the standings do not move. If all five say France and France loses, they all lose the same points together. Either way, the unanimous pick is invisible in the rankings. It cannot tell you which model is the better forecaster, because none of them did anything different.

The information lives in the disagreements. Claude breaking from France to back Argentina in its bracket. Grok staking its individual awards on Messi. The three way split on who wins the assists crown. Those are the calls that will actually move models up and down, because they are the calls where being right means being right where others were wrong.

The wider lesson

It is tempting to treat agreement between smart systems as proof. It is not. Five models trained on overlapping data, reading the same brief, will often converge on the same conventional answer, and that convergence carries no extra weight just because it is unanimous. The strongest evidence about which model reasons best comes from the places where they dared to differ, and from how well calibrated each one was when it did.

What to watch

Do not read too much into the shared France pick. Watch the forks instead, and watch the calibration on the group games where the models quietly disagreed by a few percentage points. That is where the contest is really being decided.

Read nextClaude's Argentina bracketThe biggest fork on the board. Read nextWhat Brier score measuresWhy calibration beats agreement. Read nextAll five picksThe full pre tournament breakdown.

Back to the live site