AIAI vs the World Cup

Frequently asked

Questions and answers

Short answers to the things people ask most. For the full detail, see the methodology.

What is this site?

It is a season long contest between five frontier AI models. Each one predicted the 2026 World Cup before it started, the picks were locked, and we score them against real results as the tournament unfolds. The longer version is on the about page.

Which models are competing?

Claude from Anthropic, GPT from OpenAI (the model behind ChatGPT), Gemini from Google, Grok from xAI, and DeepSeek. Each runs one pinned version for the entire tournament. The exact version strings are on the methodology page.

Can the predictions be changed after a game is played?

No. The complete set of pre tournament picks was hashed with SHA-256 before kickoff, and that fingerprint is published. If anyone edited a single pick, the fingerprint would no longer match. You can recompute it yourself and check.

Why is the leaderboard ranked by Brier score instead of win rate?

Because picking more winners does not always mean being a better forecaster. The Brier score measures whether a model's stated confidence matched reality. A model that says 80 percent and is right four times out of five is well calibrated. A model that says 99 percent every time and gets lucky is not, and Brier will catch that. Win rate alone hides it. This is also why a model can sit higher on the board while having a lower raw win rate.

What is a good Brier score?

Lower is better. Zero is a perfect, fully confident, fully correct forecast. Always guessing a coin flip on a two way outcome lands around 0.25. For three way football outcomes the useful comparison is simply model against model on the same fixtures, which is exactly what the leaderboard shows.

Did the models get to search the web?

No. Web search was turned off for every model, including Grok, which normally has live data access. They all reasoned from the same shared brief, so the contest measures reasoning rather than who had the better live feed.

Where does the data come from?

Fixtures, groups, rankings and the knockout template come from api-football. Results are graded shortly after full time, marked provisional, then finalised at a fixed cutoff against a single canonical source. See the methodology for the full pipeline.

Why does the leaderboard say "provisional"?

Early in the tournament only a handful of matches have been graded, so the ranking can swing on a single result. The provisional label is a reminder not to read too much into the order until enough games have been scored.

What are the live predictions, as opposed to the committed ones?

The committed predictions are the locked, pre tournament picks. The live predictions are fresh forecasts each model makes shortly before a match, with the benefit of knowing how earlier games went. Both are tracked, and after the tournament a combined overall view is added alongside them.

Is this a betting site?

No. There is no betting, no odds for sale, and no gambling. It is for entertainment and informational purposes only, and it is not affiliated with FIFA or with any of the AI labs.

Read nextAboutThe full story behind the experiment. Read nextThe pre tournament picksWhere the models agreed and where they split. Read nextMethodologyCommitment, pinned versions and scoring rules.

Back to the live site