How accurate is xPoints?

Before every deadline we freeze the expected-points feed. Once FPL has finalised the gameweek, we score that frozen file against what actually happened. Nothing on this page can be edited after the fact, and every number says which players it was calculated over.

Gameweeks graded
5
FPL's feed: rank match among starters
0.110
average across graded gameweeks; 1 is a perfect order, 0 is chance
FPL's feed: average error, all players
1.351
predicting zero for everyone scores 1.470
Our model minus FPL's feed: rank match
-0.051
over 2 graded gameweeks

When does our model replace FPL's feed?

not yetNot yet. The rule needs 6 graded gameweeks with our model ahead of FPL’s feed on rank match among starters, the lead clear of chance, and at least as many of the top 20 starters picked. So far: 2 of 6 gameweeks graded with both sets of numbers frozen; our model is not ahead of FPL’s feed on rank match among starters; it picks fewer of the top 20 starters than FPL’s feed.

Over the 2 gameweeks in the window our model’s rank match among starters is -0.051 against FPL’s feed (95% interval -0.183 to +0.081), and it picks 0.150 of the top 20 starters against the feed’s 0.203.

The rule is written down in the xPoints repository's score.py and evaluated over the last 6 graded gameweeks (currently gameweeks 4, 5).

Gameweek by gameweek

GWPlayers graded
all / starters
FPL's feedOur model (shadow)
Average errorZero guessRank matchTop 20Captain regretAverage errorRank matchTop 20Captain regret
GW1600 / 2101.5981.5930.086-0.05 to 0.220.1113.4n/an/an/an/a
GW2616 / 2091.4671.4500.1730.04 to 0.320.329.8n/an/an/an/a
GW3652 / 2121.2951.402-0.023-0.16 to 0.110.1013.01.100-0.044-0.16 to 0.080.086.0
GW4656 / 1971.2341.4310.2130.08 to 0.350.2515.01.0990.2300.10 to 0.350.208.0
GW5659 / 2111.1581.4730.099-0.03 to 0.240.1612.01.065-0.019-0.16 to 0.120.1011.0

Average error covers every player in the file. When the Zero guess column is red, predicting zero points for every player was closer than the feed that week. That happens because most players score zero, so we show this number but never use it to choose a model. The data behind each row (every player's prediction, minutes and points) is in the scores folder on GitHub.

Is the average right?

Expected points are an average, so the first question is whether the average comes out right. Bias is the average prediction minus the average result over every player in the file: zero is the aim, a minus sign means the forecast ran low. Squared error (root mean square) punishes big misses and is smallest for a forecast whose averages are right. The average error in the table above is kinder to a forecast that shrinks everyone towards zero, because half the players score nothing, so a lower average error is not proof of a better forecast on its own.

  • xpoints-two-stage-blend-v3 · latest gradedOver 2 gameweeks (4, 5) its average prediction ran 0.27 points a player below what happened; FPL's feed ran 0.06 points a player below what happened. Its squared error was 2.16 against the feed's 2.28, so it was closer where it counts.
  • xpoints-two-stage-blend-v2Over Gameweek 3 its average prediction ran 0.25 points a player below what happened; FPL's feed ran 0.06 points a player above what happened. Its squared error was 2.05 against the feed's 2.41, so it was closer where it counts.
GWBias, points a playerSquared errorModel version
FPL's feedOur modelFPL's feedOur model
1-0.06n/a2.62n/a
2+0.06n/a2.37n/a
3+0.06-0.252.412.05xpoints-two-stage-blend-v2
4+0.00-0.212.282.15xpoints-two-stage-blend-v3
5-0.12-0.332.292.17xpoints-two-stage-blend-v3

Over 5 graded gameweeks FPL's feed has averaged -0.01 points a player (on the mark). Bias is only meaningful over every player: among players who played, any honest forecast reads low, because it could not know who would play. Model versions are never added together; a new version starts its own record.

How to read this

Who is counted. All means every player in the frozen file. Starters means the players who played 60 minutes or more that gameweek. The same feed can look good on one group and close to random on the other. In Gameweek 1 it ordered the full list well, largely by giving near-zero to players who did not play, and only a little better than chance among those who did. So every number here says which group it covers.

Rank match is a rank correlation (Spearman). It asks whether the feed put players in the order their points actually fell. 1 is a perfect order, 0 is no better than chance. This is the number that decides whether our model replaces FPL's figure. The small range beside it is a 95% confidence interval.

Top 20 is the share of the feed's top 20 who finished in the real top 20 that week. Captain regret is the best score any player got that week minus the score of the feed's top pick, so lower is better.

Each measure has an entry in the glossary, and the methodology explains why these three and how the promotion rule uses them.

The three feeds. FPL's feed is the official expected-points figure from the game itself, and it is what the app shows today. Price only ranks players by price and nothing else, a sanity check any model must beat (it appears once the next gameweek is graded). Our model is FPL Analyst's own projection. It runs in shadow, is graded here every gameweek, and replaces FPL's figure in the app only after it has won over a run of gameweeks, not before.

How much can these numbers tell you?

One gameweek tells you very little. Rank match among starters moves by about 0.091 from one gameweek to the next through luck alone (measured from the graded gameweeks).

To be reasonably sure that one feed beats another by +0.05 in rank match takes about 26 gameweeks of evidence, or about 35 to be very sure.

Until that many gameweeks exist, a lead on this page is a hint, not proof, and the page will say so rather than round up.

What no feed can see: late team news, press conferences after the freeze, and rotation decided on the day. Predictions are frozen at the deadline, and those things happen after it.

The site's projections, by how far ahead they were made

ProjectionGameweeksAverage errorZero guessRank match, startersTop 20
FPL's feed, same gameweeks21.1961.4520.1560.20
Made at the deadline21.1961.4520.1560.20
Made one deadline earlier11.2851.4800.0710.11

At the deadline the site's average error is the same as the feed's, as it should be: the next gameweek reproduces the feed. projections made one deadline earlier carry an average error of 1.29. The feed's row is averaged over the gameweeks the deadline row covers, not the whole season. A projection made at the deadline is the feed carried through expected minutes and fixtures; the earlier it was made, the more it rests on the site's own model, so the later rows are the honest test of the planner. Every row is averaged over the gameweeks it covers; per-gameweek entries are in the site scores folder.

Expected minutes

Brier score, sixty minutes
0.081
0 is perfect, lower is better
Base-rate Brier
0.214
predicting the average rate for everyone
Skill
62%
how much better than the base rate; above zero beats it
Minutes, average error
18.8
among players who played; 13.4 over everyone
Chance of sixty minutes, as givenPlayersAverage givenActually played sixty
0% to 20%4025%5%
20% to 40%5031%30%
40% to 60%2251%59%
60% to 80%6369%75%
80% to 100%12290%94%

Over 2 graded gameweeks; the table is Gameweek 5. A well-calibrated chance means the last two columns agree in every row.

Scorecard generated 28 Sept 2026, 00:28 UTC. Latest graded gameweek: GW5. Method, source code and raw scores are on GitHub. Back to xPoints.