Behnam Analytics

Writing Data analysis & statistics

Funnel plots for comparing units of very different size

Why league tables of rates put the smallest clinics and practices at both ends, and how a funnel plot with binomial control limits and an overdispersion check fixes it.

Behnam Ebrahimi 8 min read

Rank 40 GP practices by their readmission rate and the top and bottom of the table will be crowded with the smallest practices. Not because small practices are better or worse, but because a rate built on 50 patients swings far more than one built on 1,500. A funnel plot fixes this by plotting each unit’s rate against its denominator, with control limits that are wide for small units and narrow for large ones.

This article builds one from scratch, shows what it catches and what it misses, and covers overdispersion: the reason real funnel plots often flag too many units, and what to do about it.

The example

The data is synthetic, generated with a fixed seed by the script listed at the end. It has 40 GP practices with between 34 and 1,975 discharges in a year (21,707 in total) and a 30-day readmission flag on each discharge. I planted:

  • A typical true readmission rate of 12%, with modest practice-to-practice variation, the way differences in case mix spread real practices out.
  • Practice 23: genuinely high, true rate 18%, 631 discharges.
  • Practice 31: genuinely low, true rate 8%, 966 discharges.
  • Practice 07: genuinely high, true rate 17%, but only 34 discharges.

Because the truth is known, we can check each method against it.

Why league tables mislead

Here is the top of the league table, the view a board paper usually gets:

The ten highest readmission rates, as a league table would show themDischarges in brackets
Data table
PracticeReadmission rate
Practice 30 (n=99)20.2%
Practice 23 (n=631)16.6%
Practice 20 (n=43)16.3%
Practice 11 (n=96)15.6%
Practice 22 (n=284)15.5%
Practice 13 (n=181)15.5%
Practice 19 (n=1,425)14.7%
Practice 07 (n=34)14.7%
Practice 03 (n=152)14.5%
Practice 33 (n=141)14.2%

Synthetic data. Source: projects/article-examples/stats/funnel_plots.py.

Practice 30 is “worst” at 20.2%, from 99 discharges. Its true rate in the simulation is 14.9%: above average, but not one of the planted outliers, and nowhere near 20%. Seven of the top ten practices have fewer than the median 272 discharges. At the other end it is the same story: four of the five lowest rates come from below-median practices, and one of them, Practice 15, shows 6.25% from 80 discharges while its true rate is 17.5%.

A league table has no idea how much each rate could move by chance. It treats a rate from 34 patients and one from 1,975 as equally informative, so the extremes are mostly the smallest units having a lucky or unlucky year. Whoever sits at the bottom gets asked to explain a number that mostly reflects their size.

Building the funnel

Start from the overall rate across all units, p₀. Here 11.23% of the 21,707 discharges were readmitted, so p₀ = 11.23%. If every practice had that true rate, the observed rate for a practice with n discharges would vary around p₀ with a binomial standard error of √(p₀(1 − p₀)/n). The control limits are:

  • p₀ ± 1.96 × √(p₀(1 − p₀)/n) for about 95%
  • p₀ ± 3.09 × √(p₀(1 − p₀)/n) for about 99.8%

These are the two levels most funnel plots in health use, following David Spiegelhalter’s 2005 paper Funnel plots for comparing institutional performance (Statistics in Medicine). In code:

def limits(p0: float, n: np.ndarray, z: float, phi: float = 1.0, tau2: float = 0.0) -> tuple:
    """Control limits around p0 for denominators n (normal approximation, clipped at 0 and 1).

    phi inflates the binomial variance (multiplicative overdispersion);
    tau2 adds between-unit variance (additive random effects).
    """
    se = np.sqrt(phi * p0 * (1 - p0) / n + tau2)
    return np.clip(p0 - z * se, 0, 1), np.clip(p0 + z * se, 0, 1)

Ignore phi and tau2 for now; they come in with overdispersion below. With both at their defaults, the 99.8% limits run from 0% to 25.0% for a practice with 50 discharges, 4.3% to 18.1% at 200, and 8.7% to 13.8% at 1,500. That narrowing is the funnel.

The normal approximation is fine for most denominators and is easy to write in SQL or DAX. For very small denominators, or rates close to 0% or 100%, exact binomial limits are more accurate; the lower 99.8% limit clipping to zero at small n is the tell.

Plot each practice as a point at (discharges, rate), draw the four limit lines, and add the overall rate as a reference:

30-day readmission rate by practice, against its number of dischargesBinomial control limits around the overall rate
Data table
Discharges in the yearPracticesUpper 99.8%Upper 95%Lower 95%Lower 99.8%
3414.7%28.0%21.9%0.6%0.0%
4211.9%26.3%20.8%1.7%0.0%
4316.3%26.1%20.7%1.8%0.0%
4413.6%25.9%20.6%1.9%0.0%
519.8%24.9%19.9%2.6%0.0%
5311.3%24.6%19.7%2.7%0.0%
547.4%24.5%19.7%2.8%0.0%
684.4%23.1%18.7%3.7%0.0%
806.2%22.1%18.1%4.3%0.3%
884.5%21.6%17.8%4.6%0.8%
8912.4%21.6%17.8%4.7%0.9%
9615.6%21.2%17.5%4.9%1.3%
9920.2%21.0%17.4%5.0%1.4%
10911.0%20.6%17.2%5.3%1.9%
12013.3%20.1%16.9%5.6%2.3%
1395.0%19.5%16.5%6.0%3.0%
14114.2%19.4%16.4%6.0%3.0%
15214.5%19.1%16.2%6.2%3.3%
18115.5%18.5%15.8%6.6%4.0%
26013.9%17.3%15.1%7.4%5.2%
28415.5%17.0%14.9%7.6%5.4%
28910.0%17.0%14.9%7.6%5.5%
3699.2%16.3%14.4%8.0%6.2%
37313.1%16.3%14.4%8.0%6.2%
38112.9%16.2%14.4%8.1%6.2%
40812.8%16.1%14.3%8.2%6.4%
55011.5%15.4%13.9%8.6%7.1%
61212.4%15.2%13.7%8.7%7.3%
63116.6%15.1%13.7%8.8%7.3%
72611.8%14.8%13.5%8.9%7.6%
9666.2%14.4%13.2%9.2%8.1%
103212.0%14.3%13.2%9.3%8.2%
10989.7%14.2%13.1%9.4%8.3%
129110.5%14.0%13.0%9.5%8.5%
142514.7%13.8%12.9%9.6%8.6%
154611.8%13.7%12.8%9.7%8.8%
191310.9%13.5%12.7%9.8%9.0%
192311.2%13.5%12.6%9.8%9.0%
19729.7%13.4%12.6%9.8%9.0%
19759.2%13.4%12.6%9.8%9.0%

Synthetic data. Source: projects/article-examples/stats/funnel_plots.py.

The practices that looked worst in the league table move a long way back. Practices 20 and 11 sit comfortably inside the wide mouth of the funnel. Practice 30, top of the table, is outside the 95% line but inside the 99.8% one: a warning, not a signal. Their high rates are close to what you would expect from small numbers.

Reading it: 95% and 99.8%

With 40 units and 95% limits, about two will fall outside by chance alone, even if every unit has exactly the same true rate. The 95% lines are a warning; the 99.8% lines are the ones to act on. With 40 units you would expect 0.08 of a unit outside 99.8% limits by chance, so a point out there usually means something.

In the example:

  • Outside 95%: nine practices (16, 19, 21, 22, 23, 25, 29, 30, 31). That is far more than the two you would expect by chance.
  • Outside 99.8%: three practices: 23 (high), 31 (low) and 19 (high).

Practices 23 and 31 are the planted outliers, so the funnel found them. Practice 19 is the interesting one. It has 1,425 discharges and a rate of 14.7%, and its true rate in the simulation is 13.8%. That puts it above average, but within the ordinary practice-to-practice spread I planted, not a planted outlier. The funnel flags it because, with that many patients, binomial noise alone cannot explain a 14.7% rate.

That is correct as statistics, and misleading as management. Nine practices outside the 95% limits is a symptom of the next problem.

Overdispersion

The binomial limits assume that every practice has the same true rate and that the only variation is chance. Real units are never that alike. Case mix, deprivation, coding and local services all differ, so the observed rates spread out more than binomial noise predicts. This is overdispersion, and it makes large units look like outliers for being ordinary.

Spiegelhalter’s companion paper, Handling over-dispersion of performance indicators (Quality and Safety in Health Care, 2005), sets out how to measure it and adjust for it. The measurement is short:

  1. Compute each unit’s z-score: (rate − p₀) ÷ √(p₀(1 − p₀)/n).
  2. Winsorise them: pull the top and bottom 10% in to the 10th and 90th percentiles, so the real outliers don’t inflate the estimate.
  3. The overdispersion factor φ is the mean of the squared z-scores. Unwinsorised, it sits near 1 when there’s no overdispersion. Winsorising trims the tails, so with no overdispersion at all it comes out nearer 0.7; that makes the check conservative, and a φ well above 1 is a clear sign of extra variation.
def overdispersion(z: np.ndarray) -> float:
    """phi: mean squared z-score after winsorising the top and bottom 10%."""
    lo, hi = np.quantile(z, [WINSOR, 1 - WINSOR])
    return float(np.mean(np.clip(z, lo, hi) ** 2))

Here φ = 1.58 after winsorising (3.14 before, which shows how much the outliers themselves would inflate it). To test it, compare k × φ with a chi-squared distribution: 40 × φ = 63.3, against a 95th percentile of about 54.6 for 39 degrees of freedom. The practices are more spread out than chance allows, so the limits need adjusting.

There are two ways to do it.

Multiplicative. Inflate the variance by φ, so every limit widens by √φ, here about 1.26 times. The funnel keeps its shape. In the example the same three practices stay outside the 99.8% limits, because Practice 19 is far enough out to survive a 26% widening.

Additive (random effects). Treat each practice as having its own true rate, spread around p₀ with a between-practice standard deviation τ, and add τ² to every unit’s variance:

def between_unit_variance(df: pd.DataFrame, p0: float, phi: float) -> float:
    """tau^2 for the additive random-effects limits (DerSimonian-Laird style)."""
    k = len(df)
    w = df["discharges"].to_numpy() / (p0 * (1 - p0))
    return max(0.0, (k * phi - (k - 1)) / (w.sum() - (w**2).sum() / w.sum()))

Here τ is 1.09 percentage points. The additive limits no longer shrink to nothing as n grows. At 1,500 discharges the 99.8% band becomes 7.0% to 15.4% instead of 8.7% to 13.8%. Small units barely change, because their binomial variance already dwarfs τ². With additive limits, only Practices 23 and 31 fall outside 99.8%: exactly the two planted outliers that the data can detect.

Which to use? The additive version matches the reason for overdispersion: practices really do differ, and a large practice shouldn’t be flagged for sitting within the normal spread of practices. Spiegelhalter’s paper finds both approaches reasonable but gives the additive one the stronger conceptual footing. Whichever you choose, say so on the chart, and show the unadjusted version to anyone who asks why a unit moved.

If you work in R, the NHS-R Community’s FunnelPlotR package implements these adjustments, which is useful for checking your own implementation. In Python, my SPC toolkit implements the same limits and overdispersion adjustment, and reproduces the φ and flagged practices above exactly.

What a funnel plot can’t do

  • It can’t see small units. Practice 07 really is worse, with a planted true rate of 17%, but with 34 discharges it recorded 5 readmissions (14.7%) and sits well inside the limits. Pooling several years is the usual remedy.
  • It doesn’t adjust for case mix. A practice with older, sicker patients will have more readmissions. For outcomes like mortality or readmission, plot an indirectly standardised ratio (observed ÷ expected from a risk model) against the expected count instead of a crude rate. The funnel logic is the same.
  • It doesn’t say why. A point outside the 99.8% limits is a reason to look: at the data quality first, then at case mix, then at the service.
  • It compares units at one time. To see whether a single unit has changed over time, use an SPC chart.

A checklist for your next comparison

  1. Replace ranked bars of rates with a funnel plot of rate against denominator.
  2. Draw 95% and 99.8% limits around the overall rate. Treat 95% as a warning and 99.8% as a signal.
  3. Count the units outside 95%. If it is well over 5%, estimate φ from winsorised z-scores and test it.
  4. If there is overdispersion, use additive (random-effects) limits and say so in the caption.
  5. For outcome measures, plot observed ÷ expected rather than crude rates.

Reproduce

Everything above comes from projects/article-examples/stats/funnel_plots.py, which simulates the practices, prints the league table, the flagged units under each set of limits and the overdispersion estimates, and writes both charts:

uv run python projects/article-examples/stats/funnel_plots.py