Fair to the referees, and fair to the statistics.
Refereeing is one of the few jobs assessed live, by millions, in real time, on the worst possible evidence. RefCard tries to do the opposite: assess it slowly, on the official record, with the benefit of the doubt built in.
Where the data comes from
Every grade is built on the Premier League's independent Key Match Incidents (KMI) Panel — a five-person body of former players, former coaches, and one representative each from the league and the referees' body (PGMOL). The panel reviews every major decision from every match, each week, and rules whether the on-field call and the VAR call were correct.
This matters because it means we are not substituting our own opinion for the referee's. We don't count the decisions fans were angry about. We count only the ones an independent panel formally ruled were wrong. If an incident isn't on the panel's list, we treat it as correct — even if it was controversial.
How we see the panel's verdicts. The KMI Panel doesn't publish its rulings in a public data feed. They reach the public through Dale Johnson — now BBC Sport's refereeing and VAR correspondent, previously at ESPN — whose week-by-week Key Match Incidents reviews are our source: ESPN's panel write-ups for 2023–24 and 2024–25, and Dale Johnson's BBC Sport VAR review for 2025–26. Every incident we grade is traceable to one of those published reviews, including the on-field and VAR vote split where the panel gave one.
How a grade is built
Count the panel-confirmed errors
For each referee, in each season, we take every incident the KMI Panel ruled an error — missed interventions, incorrect interventions, below-threshold mistakes, and incorrect second yellows.
Weight by severity, not by paperwork
A clear-and-obvious error the panel says VAR should have caught counts in full (1.0); a below-threshold on-field mistake that didn't meet that bar — or an incorrect second yellow — counts half (0.5). Severity drives the weight, not whether a detailed vote was published for that incident. This is why the grade doesn't track a referee's raw error count one-for-one: two officials with the same number of errors can grade differently if one's were clear-and-obvious and the other's were marginal. So a profile can read "7 errors, expected ~3" (raw counts) while the grade, which halves the below-threshold ones, sits closer to average.
Compare to the flat league average
Errors become a rate per match — so a busy referee isn't punished for volume — and are measured against a single flat baseline: what a typical official did that same season. The league averaged 0.187 errors a match in 2023–24, 0.147 in 2024–25, and 0.166 in 2025–26. We use each season's own rate rather than one pooled number, because the sourcing and the panel's standards differ from year to year. Negative means fewer errors than that baseline; positive means more.
Shrink small samples — and know when not to rank
A referee with a handful of games can look brilliant or dreadful on luck alone, so every referee is pulled toward the league average by an amount set by how much real, repeatable difference the season actually shows. When a season's referees turn out to be genuinely indistinguishable once that sampling noise is removed — as in 2025–26 — we don't rank them at all: we show the raw counts and say so.
Counts pass, rates fail — when we'll name one referee
Everything on this site is either a count or a rate, and the two carry very different weight. The test is simple enough to apply without redoing the statistics:
A count is a census. Paul Tierney worked 101 matches in the VAR booth; there were 63 confirmed on-field errors last season; Brentford were on the wrong end of 7 of them. Nobody estimated those — we counted them. There is no error bar, so there is nothing to separate, and naming the top of a count is fair.
A rate is an estimate. Fouls per game, cards per foul, errors per match, minutes added per match — each is a sample average that would come out differently with a different set of matches. It carries an error band, and two referees can only be ranked against each other if the gap between them is bigger than that band.
Almost always, it isn't. Within a single season no referee is separable from the next on any style measure — the apparent leader changes depending on how many matches you require, which means the ordering is telling you about sample sizes rather than about refereeing. Across a full 26-season career the error shrinks a great deal, and still the top few sit inside each other's range.
So we split the claim in two. We will say "this referee differs from the norm" when the difference is bigger than its own error band — over a career that clears comfortably, and it is a real, repeatable trait: split a referee's matches in half and the two halves agree closely. We will not say "this referee does it most", because that requires separating him from the referee immediately behind him, and the data does not support it.
Where you see an ordered table on this site, the order is there so you can find a referee and see roughly where he sits — not to crown the top row. We say so on the table itself.
The final number is what's left: how a referee did relative to a typical official over the same number of matches. Negative is better than the league average; positive is worse; zero is exactly average.
Why there's no difficulty adjustment
It's tempting to grade referees against the difficulty of their fixtures — surely derbies and top-six clashes breed more mistakes. We tested that, hard, and it isn't true. An earlier version of this grade did adjust for fixture difficulty; we removed it because it had no predictive power. Two independent checks agreed:
- A difficulty model built from opponent strength and match stakes, tested out-of-sample — fit on two seasons, used to predict the third — correlated with actual errors at just +0.09, indistinguishable from zero.
- A full model on all 1,140 matches, using only things known before kickoff — team quality and league position, each side's disciplinary record under other referees, derbies, matchweek, relegation and title stakes — did no better than assuming every match is equally likely to produce an error.
The finding is worth stating plainly: refereeing errors are not predictable from the character of a match. Big games and derbies are no more error-prone than a mid-table Tuesday night. So we don't pretend otherwise — every referee is measured against the same flat league average, and we'd rather say that than dress a flat number up as a difficulty model it isn't.
What the league trusts him with
Separately from the grade, each referee's profile shows the kind of fixtures PGMOL assigns him — a marker of standing and trust, not a measure of accuracy. Unlike errors, this does separate referees sharply: some are handed the marquee games week after week, others rarely. We count, pooled across seasons, how many of a referee's matches are:
- Two-big-six clashes — both clubs from a fixed list of the traditional big six: Arsenal, Chelsea, Liverpool, Manchester City, Manchester United, Tottenham. The list is fixed on purpose, so it carries no hindsight from where teams happened to finish.
- Derbies — a fixed list of recognised rivalries: Arsenal–Tottenham, Liverpool–Everton, Manchester United–Manchester City, Newcastle–Sunderland, Crystal Palace–Brighton, Aston Villa–Wolves, Chelsea–Tottenham, Chelsea–Arsenal, West Ham–Tottenham, West Ham–Chelsea, West Ham–Arsenal, and Liverpool–Manchester United.
- High-stakes run-in fixtures — matchweek 33 or later with a side in the top seven or bottom six by the table as of that round (using only earlier matchweeks — again, no hindsight).
A “marquee” match is any of the three. We report the share, the counts, and the single biggest fixture a referee took. It reads fairly at both ends: a light assignment load is a fact about the league's choices, not a mark against the official.
What we found — and won't hide
The honest headline is that Premier League referees are far more alike than the weekly outrage suggests. Most sit within a few thousandths of the average, well inside the margin of error. The real differences live at the very top and very bottom — and even those are smaller than you'd guess.
Why we built it this way
Because the alternative already exists, everywhere, and it isn't working: a referee makes a split-second call, a slow-motion replay makes it look obvious, and a verdict is reached before the restart. That cycle is unfair to officials and uninformative to everyone else. A measurement that is slow, sourced, severity-weighted, and sample-aware won't settle every argument — but it starts them from the truth.
How added time's "expected" is calculated
The added-time page compares each referee's real second-half stoppage time against what a transparent model predicts from the visible stoppages in his matches. Here is that model in full.
Go deeper
For the league-wide breakdown of what kinds of mistakes referees actually make — missed by VAR, wrongly rejected at the monitor, below the clear-and-obvious threshold — and how the totals have moved across three seasons, see the Error Anatomy page.