
On 20 July 2026, researchers Jiacheng Ding, Cong Guo, and Jason Xu published a study called WC2026-Agents, testing how well modern AI models forecast football matches, and lined the results up against the bookmaker market across the entire World Cup.
Debates over who reads a football match more accurately, algorithms or the betting market, have usually leaned on a handful of standout games. WC2026-Agents skipped the cherry-picking. The authors ran the comparison across all 104 matches of the tournament, group stage through final, rather than picking convenient examples.
How the experiment was set up
Four models went into the test: Claude Opus 4.8, ChatGPT on GPT-5.5 Thinking, Gemini 3.1 Pro, and Grok Expert Mode. Before each of the 104 World Cup 2026 matches, every model pulled fresh information on both teams. It then estimated probabilities for the 1X2 market, home win, draw, away win, before splitting a virtual $100 stake across the three outcomes, roughly the same three-step routine a human tipster would run through. After the final whistle, the same model looked back at its own call, sorting through where it had been right and where it had gone wrong.
Data collection ran a little over a month, roughly from 11 June to 19 July 2026. That produced 416 forecasts and 414 post-match reflections in all, according to the paper, each one checked against the pre-match odds bookmakers had actually set.
Who proved more accurate
All four models picked the same favorite in 92% of matches. That’s not shocking, since the models and the bookmakers are reading the same lineups, stats, and form. The real difference showed up in how each system handled the remaining 8%, the genuinely contested games. None of the four beat the bookmakers on the Brier score. That’s the sum of squared deviations between predicted probabilities and actual outcomes across all three options, and a lower number there means sharper forecasting.
Across the lines collected, the bookmaker margin averaged around 5%. That’s the built-in cushion bookmakers price into their odds, and it’s exactly what caps any edge a bettor might find: picking the right favorite mattered less than beating that markup.
The audience for Spinlander Casino’s online slots extends beyond traditional casino players, bringing together football fans, punters, statisticians and sports analysts. Interest in spinlandercasino.bet has been rising among betting market participants, particularly as more fans combine their interest in match analysis with online gaming.
A flat bet on the market favorite, placed the same way every time, outearned all four models. Their combined ROI ranged from -18% to +10%, and when the test flipped to betting against the favorite on purpose, every single model lost money. No exceptions.
The models also worked in noticeably different styles. How often a system explicitly weighed the bookmaker line in its forecast ranged from 12% to 100%, depending on the model. Admitting to a wrong call afterward was just as uneven, anywhere from 36% to 86% depending on which model you asked. That spread says more about each model’s reflective habits than about its football knowledge. The model that almost always checked itself against the odds seems to have mostly just rephrased the market’s own view, while the one that cited odds least often made more mistakes on its own, and didn’t always own up to them.
What comes next
None of the four approaches held a lasting edge over the market in money terms, and the authors say as much in their conclusions. The code and data are up on GitHub, so anyone can check the results or rerun the method on another tournament, most likely the next major European or World championship.
For the betting market, that reads less like alarm and more like confirmation of business as usual. The bookmaker line came out sharper than any of the four models, even the newest one. More people are going to ask a chatbot who’s going to win, and it’ll give them an answer, delivered with total confidence either way. Still, it’s worth checking that answer against the actual odds before getting too comfortable with how confident the algorithm sounds.


