Soccer Power Ratings: Elo, SPI, and Why They Disagree
One number per team, several systems, and built-in disagreement.
Can you really squeeze a football team — eleven players, a manager, a style, a collective mood that swings week to week — into a single number and learn anything true from it? Power ratings answer yes, and the strange thing is that they are mostly right: feed two of those numbers into a formula and you get a defensible read on who would probably win. The stranger thing, and the one this piece is really about, is that the major systems take the same matches and deliberately hand you different numbers. They are not making mistakes. They are making different bets, and once you see which bets they are placing, you can use any of them well.
How Elo works
Elo is the foundation, and it was borrowed wholesale from chess, of all places. Every team carries a rating, and the system is zero-sum: after each match, points are transferred from one side to the other. Beat a team and you take rating points from them; lose and you hand points over. Crucially, the size of the transfer depends on the surprise. Before the match, the rating gap implies an expected result — the favourite is “supposed” to win. Perform exactly as expected and ratings barely move. Pull off an upset and a large chunk of points changes hands, because the system has learned something new.
How fast it learns is set by the K-factor, a single knob governing how much any one result moves the needle. A high K makes the rating twitchy and quick to react to recent form; a low K makes it sluggish and stable, trusting the long-run history over the latest ninety minutes. Every Elo implementation has to pick a point on that responsiveness-versus-stability spectrum, and that choice alone makes two Elo systems disagree.
Football-specific Elo systems then add two adjustments the chess version never needed. The first is margin of victory: winning 4–0 is stronger evidence than scraping a 1–0, so most football Elos scale the points transfer by the goal margin rather than treating every win identically. The second is home advantage: because playing at home is worth a real and measurable edge, the expected result is computed with a bonus added to the home side before the match, so beating a strong team away earns more than the same result at home. The best-known public implementations — clubelo.com for clubs and the World Football Elo project for international sides — differ precisely in how they set K, how they weight the margin, and how big a home bump they apply. For why that home edge exists and roughly how large it is, see home advantage quantified.
Try it: turn a rating gap into a win expectancy
Put in two ratings and read the expected score straight off the curve. A 200-point edge (say 1700 vs 1500) comes out around 76%; a 400-point gap is roughly 91%; equal ratings sit at 50%.
This calculator needs JavaScript. The formula is E = 1 / (1 + 10^(−ΔR/400)), where ΔR is your team's rating minus the opponent's.
Open this calculator on the tools page →
The Soccer Power Index idea
The Soccer Power Index, or SPI, took a different angle on the same problem. Where classic Elo carries one rating per team, SPI’s central idea was to split a team into two numbers: an offensive rating — how many goals it would be expected to score against an average opponent — and a defensive rating — how many it would be expected to concede. A team’s overall strength is then derived from how those two ratings would play out against another team’s.
That separation buys real expressiveness. A devastating attack with a leaky defence and a dour side that wins 1–0 can land at a similar overall level by very different routes, and the two-number representation captures that where a single Elo figure cannot. SPI was also designed to lean on goals and chance-quality information rather than results alone, the better to estimate true strength from a limited number of matches.
One important note of historical framing: SPI was published by FiveThirtyEight, and FiveThirtyEight wound down in 2023, so its public SPI ratings and forecasts are best treated as a landmark in the history of public soccer modelling rather than a live feed you can pull today. The idea — separate offence and defence, blend goals and chance quality — remains influential and is echoed in plenty of systems still running. Treat SPI here as a design philosophy, not a current product.
Other approaches
Between pure Elo and the SPI philosophy sits a broad family of hybrid systems, and they are where most of the action is now. The defining move of the modern approach is to stop trusting results alone and feed the rating expected goals as well. The reasoning is the one that runs through this whole site: a single scoreline is a noisy, low-sample reading of how a match actually went, and a team that loses 1–0 having created the better chances has given you evidence its result hides. An xG-aware rating updates on the underlying performance, not just the final number on the board.
Providers such as Opta and others build ratings in exactly this spirit — results-plus-xG hybrids that blend who won with how comprehensively they out-created the opponent, often layering in possession-value or shot-quality signals on top. The common thread across the whole landscape is a sliding scale from results-only at one end (pure Elo: all that matters is the scoreline) to performance-aware at the other (ratings driven substantially by xG and chance quality). Where a system sits on that scale is the single biggest reason it will disagree with its neighbours, which is the subject of the next section.
Why they disagree — on purpose
Two well-built rating systems can look at an identical set of matches and rank the same teams differently, and this is a feature of their design choices rather than a bug in any of them. The disagreements cluster around a handful of decisions.
- Results-only vs. xG-aware inputs. A results-only system rewards a team for grinding out 1–0 wins; an xG-aware system may rate that same team lower if the underlying chances were poor. Same matches, different verdict.
- How they treat margin of victory. Systems that scale heavily by goal margin will rank a team that wins big above one that wins narrowly; systems that cap or ignore margin will not.
- Responsiveness (the K-factor and its analogues). A twitchy system reacts to a hot streak; a stable one trusts the body of work. They will disagree most about teams whose form has recently changed.
- Preseason priors. Where a team’s rating starts each season — carried over from last year, regressed toward the mean, or adjusted for transfers — tilts its number for weeks until results outweigh the prior.
- Transferring strength across leagues. Calibrating how a mid-table side in one division compares to a strong side in another is genuinely hard, and systems make different assumptions, so cross-league comparisons are where they diverge most.
We want to be clear that none of these choices is “wrong.” They are different bets about what best predicts the future: recent form or long-run quality, the scoreline or the performance beneath it. That is exactly the dynamic explored in why league projection models disagree — the disagreement between ratings is the disagreement between modelling philosophies, made numeric. When two reputable systems part ways on a team, the honest read is not “one is broken” but “they are weighting recent form, margin, and chance quality differently, and the truth is probably between them.”
Reading a rating gap as a win probability
A rating is only useful if you can turn the gap between two teams into something actionable, and the whole apparatus is built to let you. The core property of an Elo-style system is that the difference between two ratings maps to an expected result — the bigger your edge in rating points, the higher your win probability, along a fixed curve the system defines. A modest gap means a slight favourite and a likely close game; a large gap means a heavy favourite and a result that would barely move the ratings if it goes to form.
Two cautions make that mapping trustworthy. First, fold in the venue: because home advantage is real, the same rating gap implies a higher win probability for the home side than the away side, which is why every serious system applies a home adjustment before reading off the odds. Second, remember that football has draws and is low-scoring, so even a large rating edge translates into a meaningfully less-than-certain win probability — the favourite is favoured, not guaranteed, and a single match is a small sample. For how those match-level probabilities are constructed and how they shift as a game unfolds, see how win probability models work. Used with that humility, a power rating is what it always promised to be: not the last word on a team, but a fast, honest, single-number prior to argue with.
Sources, notes & further reading
- ClubElo — a long-running public Elo rating system for club football, with its methodology documented.
- FBref — results and xG data (via Opta) that feed performance-aware rating systems.
- Understat — season-level xG tables useful for building or sanity-checking xG-aware ratings.
- StatsBomb — research on possession value and chance quality, the inputs behind modern hybrid ratings.
Why league projection models disagree
This section was first published on 9 May 2026 as a separate article.
Two reputable forecasting models publish their Premier League title odds on the same Monday morning. One puts the leader at 72%. The other puts them at 51%. Same league, same fixtures played, same table. How? The short answer is that a projection model is a long chain of modelling decisions, and every link introduces its own error bars. This is what lives inside those chains and why small disagreements compound into very different headlines.
The basic architecture: ratings, fixtures, simulations
Almost every serious league projection follows the same three-step skeleton. First, it assigns each team a strength rating — a single number (or a pair: attack and defence) expressing how many goals they are expected to produce and concede. Second, it feeds those ratings into a match-level model: combine two teams' ratings, add a home-advantage adjustment, and you get a forecast for that fixture — typically a Poisson-distributed goal expectation for each side. Third, it simulates the remaining schedule. Each unplayed match is run thousands of times by sampling from those distributions, and the final table is recorded. Do that ten thousand times and you have a probability distribution over every possible finish: title probability, top-four probability, relegation probability.
That skeleton is largely agreed upon. What is not agreed upon is how to fill in the flesh — and those choices, quietly, are enormous.
Where the ratings come from
The most consequential decision a model makes is what evidence it uses to build team ratings. There are three broad philosophies, and they can produce meaningfully different numbers for the same team.
Results-based ratings (the Elo family) update a team's rating after every match based purely on the scoreline and the implied expectation. Beat a strong team and your rating rises; lose to a weak one and it falls. The appeal is simplicity and longevity — Elo-style systems work over any era and any data environment. The problem is that football scoreboards are noisy. A 1–0 win built on a penalty and a desperate goalline clearance counts the same as a 1–0 win built on nineteen shots and relentless possession. The model cannot tell them apart.
Expected-goals-based ratings use xG instead of actual goals: they ask not "did you score?" but "what quality of chances did you create and concede?" This strips out some of the luck in finishing and gives a cleaner read on the underlying team. The trade-off is model dependency — your rating is now downstream of your xG model, with its own assumptions about shot location, body part, and chance type. Two models using different xG sources for the same game will disagree on how strong each team is.
Market-based ratings take transfer market valuations (from databases such as Transfermarkt) as a proxy for squad quality. The appeal is that markets aggregate huge amounts of information — wages, agent negotiations, scouting networks — that a model built on shots never sees. The drawback is that valuations update slowly and encode backward-looking information; a team that has recently changed system or manager may be mispriced for months.
Many well-regarded models blend all three, weighting each source. The exact blend is a free parameter, and different teams optimise it differently.
Priors and preseason weighting
Before a ball is kicked in a new season, every team needs a starting rating. That prior is informed by the previous campaign, but how much should a twelve-month-old result affect today's forecast? This question of memory length is one of the most debated in projection methodology.
A model with a long memory treats promotion sides as roughly as bad as their third-tier record implies, and gives last season's champions significant credit even after a slow start to the new campaign. A model with a short memory updates aggressively: six strong weeks can drag a team's rating up quickly, and a bad winter can erase a title-winning reputation almost entirely.
Neither is obviously correct. Short memory is responsive; long memory is stable. The practical difference shows up most clearly mid-season: if a historically strong club is sitting eighth in November, a long-memory model may still see them as a title contender on the strength of three previous seasons, while a short-memory model has mostly moved on.
Preseason weightings also shape how signings and manager changes are treated. A model that incorporates market data can update a team's attacking rating the moment a striker is signed; a pure results-based model has to wait for evidence on the pitch. In the meantime, the two models will quote different title odds for the same club.
A worked example: how a 0.1-goal rating shift moves the needle
The relationship between a rating change and a title probability is non-linear and depends heavily on how competitive the league is, but the following illustration captures the qualitative shape.
Suppose a model rates the current leader at 1.65 expected goals per home game and 1.40 away, and their nearest rival at 1.55 and 1.30. The leader is favoured in most remaining fixtures, and across ten thousand simulations they win the title roughly six times in ten. Now suppose a second model rates the leader 0.1 goals per game lower — perhaps it is more sceptical about their underlying xG, or gives less weight to their August form. The leader is now 1.55 and 1.30: similar to the chasing pack. That gap closes their simulated title share considerably.
The intuition is that title races are decided in the close matches — the direct six-pointers, the away days at the top six. In those games, a 0.1-goal edge is a difference between being a modest favourite and being roughly evenly matched. Compounded across a season of such matches, the probability shifts are disproportionate to the rating change. This is why small methodological choices produce headline differences.
Home advantage: a number nobody agrees on
Every model adds a home-advantage term to its match predictions. Historically, playing at home in top European leagues has been worth roughly 0.35 to 0.45 additional expected goals per game — enough to shift a coin-flip fixture into a meaningful favourite-underdog structure. But that number is not fixed.
Home advantage contracted measurably during the Covid-19 period when matches were played behind closed doors, and some researchers argue it has not fully returned to pre-pandemic levels in all leagues. Whether a model uses a long-run historical average, a rolling recent estimate, or allows each club its own home-advantage value (stadiums and atmospheres vary) will shift projections, particularly for clubs whose home records diverge from the league norm.
A model that assigns a static league-wide home-advantage figure to a club with an unusually noisy home record will systematically misprice their home fixtures — and because home fixtures are roughly half the remaining schedule, that error compounds quickly.
Regression to the mean: how hard, how fast
No model takes current-season data at face value without pulling it toward a prior. A team that has won seven straight is probably good — but some fraction of that run is variance, and a model needs to decide how much of the table represents signal versus noise.
Aggressive regression treats a side on a hot streak as less dominant than the table implies; conservative regression credits the results more fully. The right answer depends on sample size: eight games into a season, regression is clearly appropriate. Forty games in, the evidence is more reliable and should be trusted further. Most models reduce the weight on preseason priors as matches accumulate, but the schedule for that reduction is itself a free parameter.
Why this matters beyond academic interest
If you follow multiple projection systems — and you should — treat the disagreements as a diagnostic tool, not a source of confusion. When two models converge on a title probability, that is relative confidence; when they diverge by twenty percentage points, one of them has made a different assumption about something meaningful. The question to ask is which assumption seems more defensible, not which number is higher.
The models worth paying attention to are the ones that publish their methodology: what ratings system, what home-advantage estimate, what regression schedule, what prior. A black-box projection that refuses to explain itself is not mysterious — it is just difficult to stress-test. Methodological transparency is the quality signal, not raw accuracy on a single season, which is dominated by luck anyway.
The disagreements between models are not a flaw — they are an accurate reflection of how much genuine uncertainty remains in a game where the best team loses roughly a third of its matches. Own the range; resist the false precision of a single number.
Where these numbers come from
- FiveThirtyEight — pioneered publicly documented club soccer projections with a published methodology; archive is a useful reference for how such systems evolved.
- FBref — xG and advanced team-level statistics across European leagues, via Opta; the raw material for xG-based ratings.
- Understat — season-long xG and xGA tables for the major European leagues.
- Transfermarkt — squad market valuations used by market-based rating systems.
- StatsBomb — detailed xG and event data; their written research covers model calibration and what event data adds over results alone.
