Methodology

Fair to the referees, and fair to the statistics.

Refereeing is one of the few jobs assessed live, by millions, in real time, on the worst possible evidence. RefCard tries to do the opposite: assess it slowly, on the official record, with the benefit of the doubt built in.

Built on
PGMOL panel
the league's independent Key Match Incidents panel — not our opinion
Seasons graded
3
2023–24 onward
A grade of 0 means
league average
the same error rate as a typical official, for the matches he actually took
Accuracy %
never computed
we count confirmed errors, not a rate — there is no complete denominator

Where the data comes from

Every grade is built on the Premier League's independent Key Match Incidents (KMI) Panel — a five-person body of former players, former coaches, and one representative each from the league and the referees' body (PGMOL). The panel reviews every major decision from every match, each week, and rules whether the on-field call and the VAR call were correct.

This matters because it means we are not substituting our own opinion for the referee's. We don't count the decisions fans were angry about. We count only the ones an independent panel formally ruled were wrong. If an incident isn't on the panel's list, we treat it as correct — even if it was controversial.

How we see the panel's verdicts. The KMI Panel doesn't publish its rulings in a public data feed. They reach the public through Dale Johnson — now BBC Sport's refereeing and VAR correspondent, previously at ESPN — whose week-by-week Key Match Incidents reviews are our source: ESPN's panel write-ups for 2023–24 and 2024–25, and Dale Johnson's BBC Sport VAR review for 2025–26. Every incident we grade is traceable to one of those published reviews, including the on-field and VAR vote split where the panel gave one.

How a grade is built

Count the panel-confirmed errors

For each referee, in each season, we take every incident the KMI Panel ruled an error — missed interventions, incorrect interventions, below-threshold mistakes, and incorrect second yellows.

Weight by severity, not by paperwork

A clear-and-obvious error the panel says VAR should have caught counts in full (1.0); a below-threshold on-field mistake that didn't meet that bar — or an incorrect second yellow — counts half (0.5). Severity drives the weight, not whether a detailed vote was published for that incident. This is why the grade doesn't track a referee's raw error count one-for-one: two officials with the same number of errors can grade differently if one's were clear-and-obvious and the other's were marginal. So a profile can read "7 errors, expected ~3" (raw counts) while the grade, which halves the below-threshold ones, sits closer to average.

Compare to the flat league average

Errors become a rate per match — so a busy referee isn't punished for volume — and are measured against a single flat baseline: what a typical official did that same season. The league averaged 0.187 errors a match in 2023–24, 0.147 in 2024–25, and 0.166 in 2025–26. We use each season's own rate rather than one pooled number, because the sourcing and the panel's standards differ from year to year. Negative means fewer errors than that baseline; positive means more.

Shrink small samples — and know when not to rank

A referee with a handful of games can look brilliant or dreadful on luck alone, so every referee is pulled toward the league average by an amount set by how much real, repeatable difference the season actually shows. When a season's referees turn out to be genuinely indistinguishable once that sampling noise is removed — as in 2025–26 — we don't rank them at all: we show the raw counts and say so.

Counts pass, rates fail — when we'll name one referee

Everything on this site is either a count or a rate, and the two carry very different weight. The test is simple enough to apply without redoing the statistics:

A count is a census. Paul Tierney worked 101 matches in the VAR booth; there were 63 confirmed on-field errors last season; Brentford were on the wrong end of 7 of them. Nobody estimated those — we counted them. There is no error bar, so there is nothing to separate, and naming the top of a count is fair.

A rate is an estimate. Fouls per game, cards per foul, errors per match, minutes added per match — each is a sample average that would come out differently with a different set of matches. It carries an error band, and two referees can only be ranked against each other if the gap between them is bigger than that band.

Almost always, it isn't. Within a single season no referee is separable from the next on any style measure — the apparent leader changes depending on how many matches you require, which means the ordering is telling you about sample sizes rather than about refereeing. Across a full 26-season career the error shrinks a great deal, and still the top few sit inside each other's range.

So we split the claim in two. We will say "this referee differs from the norm" when the difference is bigger than its own error band — over a career that clears comfortably, and it is a real, repeatable trait: split a referee's matches in half and the two halves agree closely. We will not say "this referee does it most", because that requires separating him from the referee immediately behind him, and the data does not support it.

Where you see an ordered table on this site, the order is there so you can find a referee and see roughly where he sits — not to crown the top row. We say so on the table itself.

The final number is what's left: how a referee did relative to a typical official over the same number of matches. Negative is better than the league average; positive is worse; zero is exactly average.

Why there's no difficulty adjustment

It's tempting to grade referees against the difficulty of their fixtures — surely derbies and top-six clashes breed more mistakes. We tested that, hard, and it isn't true. An earlier version of this grade did adjust for fixture difficulty; we removed it because it had no predictive power. Two independent checks agreed:

The finding is worth stating plainly: refereeing errors are not predictable from the character of a match. Big games and derbies are no more error-prone than a mid-table Tuesday night. So we don't pretend otherwise — every referee is measured against the same flat league average, and we'd rather say that than dress a flat number up as a difficulty model it isn't.

What the league trusts him with

Separately from the grade, each referee's profile shows the kind of fixtures PGMOL assigns him — a marker of standing and trust, not a measure of accuracy. Unlike errors, this does separate referees sharply: some are handed the marquee games week after week, others rarely. We count, pooled across seasons, how many of a referee's matches are:

A “marquee” match is any of the three. We report the share, the counts, and the single biggest fixture a referee took. It reads fairly at both ends: a light assignment load is a fact about the league's choices, not a mark against the official.

What we found — and won't hide

The honest headline is that Premier League referees are far more alike than the weekly outrage suggests. Most sit within a few thousandths of the average, well inside the margin of error. The real differences live at the very top and very bottom — and even those are smaller than you'd guess.

Where we're honest about the limits. A single season is a noisy read — year-to-year stability is only moderate, so a one-season swing is usually variance, not a real change in ability. That's why we show three-season trajectories, not just a snapshot.
On the 2023–24 data. For that season, the official record gives us confirmed errors but not the detailed per-incident votes. We weight those errors by their severity rather than penalising them for missing paperwork — and we flag the season's grade as built on a slightly coarser source than the two that follow.
On small samples. Referees with fewer than 15 matches in a season are shrunk hard toward the average and shown separately. Their grades are estimates with wide uncertainty, not verdicts.
On 2025–26. Once sampling noise is removed, this season's referees are statistically indistinguishable on the severity-weighted grade — the whole field sits within a few thousandths. So we don't rank it. We still show each referee's raw errors against the league-average baseline, because the counts are real; we just don't turn a difference that small into a league table.

Why we built it this way

Because the alternative already exists, everywhere, and it isn't working: a referee makes a split-second call, a slow-motion replay makes it look obvious, and a verdict is reached before the restart. That cycle is unfair to officials and uninformative to everyone else. A measurement that is slow, sourced, severity-weighted, and sample-aware won't settle every argument — but it starts them from the truth.

How added time's "expected" is calculated

The added-time page compares each referee's real second-half stoppage time against what a transparent model predicts from the visible stoppages in his matches. Here is that model in full.

explains of added-time variance (R²)
Read these as averages, not measurements. Each number is how much added time tends to rise with one more of that event, across 658 matches. The VAR figure is the clearest example: we count VAR decisions, we don't time them — a 20-second check and a four-minute review both count as one. So "VAR ≈ +1.5 min" is an association, not the length of a check. Injuries and actual VAR-check durations aren't in the data, which is why the model leaves most of the variance unexplained.

Go deeper

For the league-wide breakdown of what kinds of mistakes referees actually make — missed by VAR, wrongly rejected at the monitor, below the clear-and-obvious threshold — and how the totals have moved across three seasons, see the Error Anatomy page.