El ColegiadoData on refereeing in the Liga

Analysis

Methodology

How every figure on the site is produced, and what it cannot tell us.

  • 2000-01 → 2025-26
  • 9,880 matches
  • Updated on 3 October 2026

Coverage: Liga only (Primera División). Copa del Rey and Champions League: coming later.

Sources

  • football-data.co.uk: scores and pre-match odds from 2000-01 to 2025-26; fouls, corners, shots and cards from 2005-06. Free to use; the download is done by hand (the site's robots.txt excludes automated agents).
  • Zenodo 7341037 (BeSoccer data): referee, cards and penalties per team, up to 2021-22. Licence CC BY-NC 4.0: non-commercial project, attribution required.
  • Zenodo 7831873 ("Dataset with the results and matches of La Liga"): the referee of every match from 2014-15 to 2022-23 (up to matchday 29 of 2022-23). Licence CC BY-SA 4.0.
  • Wikipedia (en and es): referees and penalties of each match taken from every club's season articles, final standings, champions and referee biographies. Licence CC BY-SA 4.0; each reference is dated by its revision number.
  • Referee retained: up to three sources (Zenodo/BeSoccer, Zenodo 7831873, Wikipedia) vote; the majority wins, and a disagreement without a majority leaves the referee blank (conflict log on the Data page).

Sources set aside: sites whose terms of use forbid collection, or that are protected against bots (no protection was circumvented).

Standardisation

Teams: an alias table links every spelling to a single identifier; an unknown name makes the build fail.

Referees: identifier = the two surnames without accents. The spelling variants encountered are listed on each profile. 80 referees are identified.

The Teixeira Vitienes brothers: Fernando Teixeira Vitienes and José Antonio Teixeira Vitienes both refereed in Primera. The main source conflates them from 2010-11; when no other source gives the first name, the match is filed under a separate, undistinguished entry (44 matches), excluded from the referee × club tests.

Penalties awarded = penalties scored + penalties missed by the team. Odds: market average, otherwise the average of the available bookmakers.

Periods

  • Before: 2000-01. According to the public prosecutor the payments began "at least in 2001": this single season serves as a sensitivity check.
  • During: 2001-02 → 2017-18, the period of the payments.
  • After: 2018-19 → 2025-26. Coincides with the arrival of VAR (2018-19): the two effects cannot be separated.

Odds: which figure, which correction

  • Source: football-data.co.uk, pre-match market average (columns BbAv* from 2005 to 2019, Avg* from 2019 to 2026; before 2005, or when no average is available, the average of the available bookmakers). These are not closing odds.
  • Margin removed by proportional normalisation: p = (1/odds) / Σ(1/odds) over the three outcomes. Shin's method is not used; proportional normalisation may leave a slight favourite–longshot bias (see limitations).
  • 2 matches out of 9,880 without odds: excluded from the "points − expected" measures and from the strength adjustments, kept everywhere else.

Expected points

For each match, the probabilities of a win, draw and defeat derived from the odds give the expected points.

Expected points = 3 × P(win) + P(draw)

The deviation points won − expected points neutralises the team's strength: a big club has many expected points, so it has no "automatic" positive deviation. Across the whole Liga, this deviation ranges from −1.3% to +1.8% depending on the season: the odds are a reliable benchmark (see the study).

Exact distribution of points

The site's signature chart computes, in your browser, the exact probability of every points total over a series of matches if only the odds mattered. It starts from a "0 points with certainty" distribution, then, match after match, combines it with the three possible outcomes: 0 points (defeat), 1 point (draw), 3 points (win), each with its probability. No random draws: the result is exact.

The two-sided p-value of a discrete distribution is defined the same way everywhere:

p = 2 × min( P(X ≤ observed), P(X ≥ observed) ), capped at 1

In the analysis (Python), the probabilities come from a simulation, with the usual "+1" correction ((k + 1) / (n + 1)); on the site they are computed exactly. The smaller p is, the rarer the actual total is in light of the odds. In the Explorer and the comparison tool, p is not corrected for multiple comparisons: it is expressed as "common" (p ≥ 0.10), "uncommon" (0.01 ≤ p < 0.10) or "rare" (p < 0.01).

Referee × club tests

Every referee × club pair with at least 5 matches is tested on five measures (1,325 pairs, all clubs):

  • Points − expected: the series of matches is replayed thousands of times, drawing each result according to the odds; the p-value is how often chance alone produces a deviation at least as large.
  • Balances of penalties, red cards, yellow cards and fouls (stratified permutation): the pair's matches are compared with the same club's other matches in the same season under other referees, by randomly reallocating matches between referees.

Below 15 matches, the pair is flagged "small sample". Below 5 matches, it is not tested: "descriptive" chip. A dedicated page exists from 3 matches.

Benjamini-Hochberg correction

With hundreds of comparisons, about 5% give p < 0.05 by pure chance. The Benjamini-Hochberg correction ranks the p-values and adjusts them to the number of tests, giving a q-value. A deviation is only called notable ("favourable to the club" or "unfavourable to the club") if q < 0.05; this keeps the expected share of false alarms below 5%. Otherwise: "inconclusive".

The families of tests corrected together:

  • Referee × club: one family per measure (points, penalties, red cards, yellow cards, fouls), each over all pairs from all clubs with at least 5 matches (1,325 pairs; slightly fewer for fouls, unavailable before 2005-06).
  • Appointments: one family (2,094 pairs).
  • Important matches: one family (1,297 combinations).
  • The study's differences-in-differences: one family (18 tests).

The count of "notable results" adds up the q < 0.05 across all referee × club families.

Why so many "inconclusive" results?

Strength adjustment

A dominant club wins more penalties and receives fewer cards because it dominates, not necessarily because of the referee. Each match is therefore compared with the Liga average for matches of the same strength according to the odds (deciles), in the same season and at the same venue (home or away). A positive value is favourable to the club.

Difference-in-differences

We measure how a club changes between two periods, then subtract the change of a comparison group over the same periods: (A₁ − A₂) − (B₁ − B₂). The 95% confidence interval comes from a club-season block bootstrap (whole club seasons are resampled, to respect the dependence between matches of the same season). In the study, q = Benjamini-Hochberg correction over the 18 tests. The comparison tool repeats this calculation in the browser for any club and any control group, with 2,000 draws and a fixed seed.

Appointment frequency

Expected number of matches of a referee with a club = his share of Liga matches each season (among the matches he may officiate: never a club from his own region) × the club's matches. Poisson test, Benjamini-Hochberg correction.

Important matches

A match is "important" for a club if the opponent finished in the season's final top 4 (recalculated standings: points, head-to-head, goal difference, goals scored), plus the Clásico for Barça and Real Madrid. Period: 2000-01 → 2025-26 (referees known for all matches). For each club and each referee (at least 5 matches), the share of important matches he officiated is compared with that of the other referees: Fisher's exact test, then Benjamini-Hochberg correction.

Archives 1990-91 → 1999-00

Sources: results and final standings from es.wikipedia season pages (complete); referee, date, penalties and cards from the clubs' season pages on en.wikipedia, when they provide them (CC BY-SA 4.0).

Coverage: referee known for 1,112 matches out of 3,964; this coverage is partial and depends on the clubs and on Wikipedia contributors, hence not random. Two points for a win until the season before 1995-96.

Referee names: cautious consolidation of spelling variants: when in doubt, two spellings are kept apart rather than wrongly merged; a few duplicates may therefore remain.

Why no analysis: no bookmaker odds exist for these seasons, so there are no expected points, and the refereeing data are incomplete and not random. An Elo model (expected points computed from past results only) was built and tested against the odds over 2000-01 → 2025-26; its expected points differ by several points per club and per season from those derived from the odds, so it was rejected as a benchmark (see scripts/09_elo.py). The archive pages are therefore purely descriptive: no tests, no verdicts.

Limitations

  • Odds reflect what the market knew: if a bias were known to bookmakers and priced in, the deviation from the odds would underestimate it.
  • Pre-match market average rather than closing odds, margin removed proportionally (no Shin method): a slight favourite–longshot bias is possible.
  • Liga only: the Copa del Rey and European competitions are not covered.
  • Cards and fouls depend on playing style (possession, pressing), which the odds-based adjustment does not capture.
  • The end of the payments and VAR coincide (2018-19); matches behind closed doors in 2019-20 and 2020-21.
  • Yellow cards: the two sources agree on 85 to 88% of matches; a single-source series serves as a check.
  • The final standings shown are recalculated from the matches (points, head-to-head, goal difference, goals scored); the official title prevails.
  • These tests concern aggregate results. They say nothing about any specific decision in a match, and a statistical deviation is not proof.