World Cup 2026

Simulating the Knockout Rounds: Why Favourites Lose More Than You'd Think

A favourite that wins 70% of its matches still clears five rounds only about one time in six. Here is the arithmetic.

Ask someone to name the best team left in a knockout bracket and they will answer with conviction. Ask them how often that team actually lifts the trophy and the conviction wobbles, because the honest answer is "less than you think." A single-elimination tournament is a machine for converting small per-match edges into large doses of randomness, and the only way to feel that properly is to simulate it — to play the bracket out thousands of times and count the endings. This piece is about how that simulation works, why it keeps embarrassing the favourite, and where its answers stop being trustworthy. You can run one yourself at the bottom of the page; the World Cup 2026 simulator is embedded inline so you can watch a top side fail to win as often as your gut insists it should.

What a knockout simulation actually is

Strip away the interface and a tournament simulation is a loop. You give every team a strength number. You write down a rule that turns two strength numbers into a match result — usually a goal model, where the rating gap sets how many goals each side scores on average and the actual scoreline is drawn at random around those means. Then you play the whole bracket once: every tie resolved by the goal model, the winner advancing, until one team is left holding the trophy. That is a single simulated tournament. Do it once and you have learned almost nothing — it is one roll of a very large set of dice. Do it ten thousand times and tally how often each team reached each round, and those tallies become probabilities: the share of simulated tournaments a team won is its title odds under your assumptions.

The phrase "under your assumptions" is doing heavy lifting, and we'll come back to it. But the core machinery is genuinely that simple. There is no oracle inside, no secret knowledge of who is in form — just a strength rating, a match model, and a very patient loop. The methodology in prose, with the modelling choices spelled out, lives in how a World Cup simulation works; here we want to dwell on the result that surprises people most.

The loop in one breath
Rate the teams → turn each rating gap into a match result with a goal model → play the bracket to a champion → record who reached each round → repeat thousands of times → the counts are the odds. The simulation invents nothing; it just replays your assumptions until the randomness averages out.

Why favourites lose: the arithmetic of compounding

Here is the heart of it. Suppose you have a genuinely excellent team — strong enough that, in any single match against the kind of opponent it will face in a deep bracket, it wins 70% of the time. That is a big edge; most real favourites are nowhere near that dominant against fellow survivors. Now ask the question that matters: how often does this team win five straight matches to take the trophy?

The answer is not 70%. It is 0.70 multiplied by itself five times, because the team has to win the first match and the second and the third and the fourth and the fifth. That product is 0.705 ≈ 0.168 — about one tournament in six. A team you would describe, with total justification, as the clear best side in the field still fails to win the thing more than 80% of the time. Drop the per-match win rate to a still-excellent 60% and the title probability collapses to 0.605 ≈ 0.078, fewer than one run in twelve. The edge that feels enormous over a single match is eaten alive by the requirement to repeat it.

Compounding upset odds
P(win the tournament) = (per-match win rate)number of rounds, when each round is roughly independent.
70% per match over 5 rounds → 0.705 ≈ 16.8%.
60% per match over 5 rounds → 0.605 ≈ 7.8%.
The favourite's edge survives one match; it does not survive five.

Flip the same arithmetic around and you see why upsets feel so common at tournaments. The probability that the favourite does not win is 1 minus its title chance — better than 83% in the 70%-per-match case. Somebody other than the best team wins most of the time, not because the best team is overrated, but because single-elimination is a format that hands the field, collectively, more chances than it hands the single strongest entrant. Spread a little win probability across thirty-one other teams and their combined claim on the trophy dwarfs the favourite's. This is the structural reason the favourite wins less often than its quality suggests, and it is baked into the bracket before a ball is kicked.

One big chance, replayed: a hypothetical to feel the variance

Numbers on a page are bloodless, so picture a clearly-labelled hypothetical. Two teams meet in a quarter-final. Team A is the better side — say the goal model gives it a 1.6 expected-goals night against Team B's 1.1. Feed those into the kind of scoreline distribution the expected-points model uses and Team A is the favourite, but only at roughly a coin-flip-and-a-bit: it wins maybe a little under half the time outright, with a big slice of draws that then go to penalties, where its edge nearly vanishes. Over a league season, A finishes comfortably above B and nobody argues. Over this one match, B goes through often enough that when it happens, it barely qualifies as a shock. (Those xG figures are an illustrative construction to show the mechanism — not a real or predicted fixture.)

Now chain four of those nights together. Even if A were a slight favourite in each, the chain of "win, then win, then win again" is where the favourite's hold loosens. The simulation does this chaining for you, thousands of times, and the title-odds column it spits out is just the honest accounting of how rarely the chain completes. That is the whole value of running it rather than reasoning about a single tie: a human brain anchors on "A is better, so A goes through," while the simulation patiently reminds you that "better in each match" and "wins the bracket" are very different claims separated by an exponent.

Run it yourself

The simulator below plays the real 48-team, 12-group World Cup 2026 format — twelve groups of four, top two plus the eight best third-placed teams into a 32-team knockout — over as many Monte-Carlo tournaments as you like. Crucially, it ships with generic placeholder teams and ratings that you overwrite: it knows nothing about the actual 2026 results, draw, or standings, and it never will. Set your own strengths, run a few thousand tournaments, and watch the title-odds column. How big the top number gets depends almost entirely on how far clear of the field you put your best side: leave the shipped placeholder spread alone, where the ratings step down gently from 90 to 38, and even the top-rated team takes fewer than one title in twenty. Nudge that rating up five points and re-run: the share barely twitches. Push every rating to the extremes the inputs allow, 100 down to nearly zero, and it still only reaches about 7% — roughly one tournament in fourteen for a side rated further clear of the field than any real team has ever been. That stubbornness is the variance of the format talking, and it is the whole reason the simulation is worth running.

This simulator needs JavaScript. The method: simulate each group as a round robin with Poisson goals (scoring rate set by the rating gap), rank by points then goal difference, advance the top two of each group plus the eight best third-placed teams into a 32-team knockout decided by the same goal model, and repeat thousands of times to estimate each team's advance and title odds. It uses placeholder teams you edit — no real 2026 results are encoded.

Open the full simulator on its own page →

What the simulation tells you

Used honestly, a knockout simulation answers a specific and useful set of questions. It tells you the shape of the variance: how flat the title race is, how much of the field has a realistic claim, how steeply a team's chances fall as it has to survive more rounds. It tells you the sensitivity of the outcome to your inputs — change a rating and see whether the odds lurch or barely flinch, which reveals how much the bracket is decided by strength versus luck. It gives you a baseline against which to judge a real result: if a side you rated middling reaches a semi-final, the simulation tells you whether that was a one-in-three fluke or a one-in-fifty miracle. And it makes the compounding visceral in a way no single probability does, because you watch favourite after favourite fall out across the runs.

What it does not tell you

The cautions matter as much as the capabilities, and a simulation that you trust uncritically is worse than no simulation at all. First, it is only as good as the ratings you feed it — garbage strengths in, confident-looking garbage out. The crisp percentages tempt you to forget they rest on a number you essentially guessed. Second, it assumes a stable model of a match: it has no idea about a key suspension, a tactical mismatch, a team that raises its level for big games, fatigue from a brutal travel schedule, or the simple fact that knockout football is often cagier and lower-scoring than the group stage. Third, independence between rounds is an approximation — a team that grinds through extra time twice may be more vulnerable in the next round than a fresh opponent, and the basic model misses that. Fourth, and most important here, it is emphatically not a forecast of the real tournament: it is a property of your assumptions and the format's structure, a teaching instrument, not a prediction of or commentary on actual matches. The deeper warning against reading too much into any single bracket — simulated or real — is laid out in don't overfit the knockouts.

How to read a simulation well

Treat the title-odds column as a distribution, not a verdict. The useful read is rarely "this team will win" — it is "no single team is close to a favourite," or "my top two ratings barely separate at the top," or "the field is so deep that the trophy is genuinely up for grabs." Run the thing multiple times with deliberately different ratings to see which conclusions are robust to your uncertainty and which evaporate the moment you admit you might be wrong about a team. Watch how the per-stage odds decay round by round, because that decay is the compounding, made into a column you can point at. And hold the whole exercise at arm's length from the real World Cup: the simulator models the mathematics of a 48-team bracket beautifully and tells you nothing whatsoever about who is actually winning in 2026. Used that way — as a lens on variance rather than a window onto the future — it is one of the most honest tools in the analytics kit, precisely because it keeps refusing to crown your favourite as often as you would like.

Sources, notes & further reading

How a World Cup simulation works

This section was first published on 5 June 2026 as a separate article.

When a forecaster says a team has, say, a one-in-eight chance of winning the World Cup, they did not derive it from an equation. They built a model of the tournament, played the entire thing from group stage to final, recorded who won, and then did it again tens of thousands of times. The team's title probability is simply the fraction of those simulated tournaments it won. This is the Monte Carlo method, it is how essentially every serious World Cup forecast is produced, and understanding it tells you exactly why two careful models can disagree about the favourite.

Why you cannot just calculate it

It is tempting to think title odds should come from a formula — multiply the probability of winning the group by the probability of winning each knockout round. The problem is that those probabilities are not independent or fixed. Who a team meets in the round of 16 depends on which teams top their groups, which depends on results that have not happened; a kind path can open or a brutal one can close based on a single upset three games earlier. The bracket is a tree of contingencies, and the clean way to handle a tree of contingencies is not to solve it analytically but to sample it: play it out, let the contingencies resolve themselves, and repeat until the averages stabilise.

The two ingredients

A simulation needs exactly two things. The first is a rating for every team — a strength estimate, whether a single Elo figure or a paired attack-and-defence number in the SPI tradition. Where those come from, and why reasonable people build them differently, is covered in soccer power ratings: Elo, SPI and why they disagree and how models rate the field before a World Cup is played.

The second is a match model that turns two ratings into the result of a single game. The workhorse is a Poisson model: convert the rating gap (plus any venue adjustment) into an expected number of goals for each side, then treat each team's goals as a draw from a Poisson distribution. That yields a probability for every scoreline, and therefore for a win, draw or loss. We build precisely this engine, step by step, in the Poisson goals model in Python, and explain its place in forecasting in how win-probability models work.

One simulated tournament
For every group game, draw a scoreline from the match model and bank the points. Rank each group, advance the qualifiers, then for each knockout tie draw a result (settling ties by a near-coin-flip shoot-out) until one team remains. Record the winner. That is one run; a forecast is fifty thousand of them.

Walking through a single run

One simulated tournament proceeds exactly as the real one would. Each group fixture is played by sampling a scoreline from the match model, and the points are tallied. The group tables are sorted — with tiebreakers applied as the competition rules specify — and the qualifying teams advance into the knockout bracket. Each knockout tie is then sampled in turn; because knockouts cannot end level, a tie that the model lands on as a draw is pushed to extra time and, if needed, a shoot-out, which most simulations treat as close to 50/50 between two teams of any strength. Eventually one team is left standing, and the simulation notes the champion, the finalists, the semi-finalists, and so on.

That single run is almost meaningless on its own — it is one possible tournament, dominated by the luck of the draws it happened to sample. The power comes from repetition. Run it 50,000 times and the noise averages out: a team that won 6,000 of those runs gets a 12% title probability, and every other question — reach the semis, top the group, suffer a group-stage exit — is answered the same way, by counting the fraction of simulations in which it occurred.

Why the number is a range, not a fact

Two competent modellers can run this identical procedure and publish different favourites. The disagreement does not come from the simulation engine — that part is mechanical — but from the inputs, and a handful of choices dominate.

The ratings. The biggest lever. A results-only Elo and a chance-quality-aware hybrid can rate the same team meaningfully differently, and because the match model is non-linear, a modest rating gap can swing a title share by several points. This is the same compounding sensitivity that makes club projections diverge, dissected in why league projection models disagree.

The home and venue edge. How much advantage to give host or near-host teams, and how to handle travel, altitude and heat across a continent-spanning tournament, is a real and unsettled choice — see altitude and heat at the 2026 venues and the general magnitude in home advantage, quantified. A model that bakes in a large host edge will tilt its whole distribution toward those teams.

The shoot-out assumption. Treating shoot-outs as pure coin flips versus giving a small edge to the stronger side, or to teams with documented preparation, nudges the deep-run odds of every contender. The case that shoot-outs are partly trainable is made in preparing for penalties.

Variance settings. How much randomness the match model injects controls how often upsets happen. A higher-variance model produces more chaos and flatter odds; a lower-variance one concentrates probability on the favourites. Neither is obviously right, and the choice shows up directly in how open the field looks.

Reading a simulation honestly

The output of a Monte Carlo run is genuinely useful — it is a coherent, internally consistent map of how a tournament could unfold, and it correctly captures things human intuition mangles, like the value of a kind draw or the brutal arithmetic of needing seven straight results. But it is a map drawn from assumptions, not a measurement of the future. The right way to use two such forecasts is the way we recommend for league models: when they agree, that is relative confidence; when they diverge, one of them has made a different, examinable assumption about ratings, home edge, or variance. The honest takeaway is to read the spread across models as the real uncertainty, and to distrust any single decimal that claims to know the winner of a tournament that, by its nature, the best team loses more often than not. For how all of this frames the expanded field, start with World Cup 2026 by the numbers.

Sources & further reading

  • StatsBomb — event data and research underpinning chance-quality ratings and match models.
  • FBref — international results and xG, the inputs a simulation's ratings are trained on.
  • ClubElo — a transparent rating system whose published method shows the kind of choices that drive simulation disagreement.
  • FIFA — official tournament format, group structure and knockout bracket rules that a simulation must encode.