Writing Data analysis & statistics
A/B testing in operations
How to run a fair experiment on a reminder text, a booking letter or a clinic template, from the randomisation unit and the sample size to intention to treat, confidence intervals and governance sign-off.
Most operational changes are judged with a before-and-after chart. The new reminder text went live in March, the DNA rate was lower in April, so the text worked. But April also had different clinics, different weather, a different mix of patients, and whatever else changed that month. A before-and-after comparison can’t separate any of that from the change itself.
A randomised experiment can. Some patients (or clinics, or days) get the new version and some don’t, chosen by chance, at the same time. Anything else that happens that month hits both groups equally, so the difference between them is the effect of the change. Online companies call this A/B testing. The method is the same for a reminder text, a booking letter or a clinic template, and so are the decisions that make or break it.
Why before-and-after misleads
In my appointment reminder experiment, the pre-trial DNA rate was 10.0%. During the trial, the arm that got no reminder at all ran at 9.3%. A before-and-after comparison would have credited the texts with that 0.7-point drop too, and it had nothing to do with them. The data is synthetic, but real services see the same thing: the rate drifts, changes often go in after a bad month (so some improvement would have come anyway), and two initiatives launch in the same quarter.
An SPC chart can tell you that a process changed and when (SPC charts for operational metrics covers how). It can’t tell you why. When the question is “did our change cause this?”, randomise.
Choose the unit of randomisation
The unit is whatever gets allocated to one version or the other. Pick it to match how the change is delivered and where it could leak between groups.
| Change | Randomise by | Why |
|---|---|---|
| Reminder text wording | Patient | Randomising appointments would send one patient both versions |
| Booking letter design | Patient or referral | The letter goes to one person about one pathway |
| Clinic template (slot length, overbooking) | Clinic or clinic session | Every patient in a session shares the template |
| Call-handling script | Call handler, or day | A handler can’t switch scripts between calls cleanly |
Two consequences follow.
Analyse at the level you randomised, or cluster. If patients are randomised but you count appointments, a patient’s repeat appointments are correlated, and a standard test treats them as independent. The intervals come out too narrow. In the project, the “95%” intervals that ignored this covered the true effect only 92% to 93% of the time. Cluster the standard errors by patient, or analyse one row per patient.
Clusters cost sample size. The design effect, 1 + (m − 1) × ICC, says how much the variance grows, where m is the number of observations per unit and the ICC (intraclass correlation) is how alike observations in the same unit are. A patient with 2.4 appointments and an ICC of 0.1 gives a design effect of 1.14: 14% more appointments needed. Randomise whole clinics with 200 appointments each, and an ICC of just 0.01 gives 2.99; at 0.02 it’s 4.98. Clinic-level trials need many more appointments, and enough clinics that chance doesn’t hand one arm all the busy ones. Pairing similar clinics before randomising helps.
Whatever the unit, generate the allocation in code with a fixed seed, save it before the change goes live, and make sure the people booking can’t see or influence who gets what. “Odd NHS numbers get the new letter” is predictable, and anything predictable can be steered.
Sample size and power for a rate
A sample size calculation for a rate needs four numbers:
- The baseline rate, from recent data: say a 10% DNA rate.
- The minimum detectable effect (MDE): the smallest change worth finding. A 2-point fall, to 8%, might be the smallest a service would act on.
- The significance level: the false-positive rate you’ll accept, usually 5%, two-sided.
- The power: the chance of detecting the MDE if it’s real, usually 80% or 90%.
statsmodels does the calculation. proportion_effectsize converts two rates into Cohen’s h, 2 * (arcsin(sqrt(p1)) - arcsin(sqrt(p2))), and NormalIndPower().solve_power returns the sample size of the first group, with the second set by ratio. This is the function the project uses:
import math
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize
def appointments_per_arm(
baseline: float, mde: float, alpha: float, power: float, design_effect: float = 1.0
) -> int:
"""Appointments needed in each arm to detect a fall of `mde` in a DNA rate of `baseline`."""
effect = proportion_effectsize(baseline, baseline - mde)
n = NormalIndPower().solve_power(effect_size=effect, alpha=alpha, power=power, ratio=1.0)
return math.ceil(n * design_effect)
appointments_per_arm(0.10, 0.02, 0.05, 0.80) returns 3,205. As a cross-check, power_proportions_2indep(-0.02, 0.10, 3205, alpha=0.05).power from the same module, which uses a pooled variance under the null and separate variances under the alternative, gives a power of 0.799 for that size. The two methods agree.
How the answer moves with the inputs, for a 10% baseline:
| Fall to detect | 5% level, 80% power | 5% level, 90% power | 1.67% level, 80% power |
|---|---|---|---|
| 1 point | 13,488 | 18,056 | 17,990 |
| 2 points | 3,205 | 4,291 | 4,275 |
| 3 points | 1,347 | 1,803 | 1,796 |
Appointments per arm. Halving the effect you want to detect roughly quadruples the sample. The last column is for three comparisons against one control, with the 5% split three ways (Bonferroni) so that the chance of any false positive stays at or below 5%. Every extra arm costs sample size twice: more groups to fill, and a stricter threshold for each.
Turn it round: what can you detect?
Operational trials often have a fixed window: the appointments that will happen in the next 12 weeks. Then the useful question is the reverse one. Given this many appointments, what’s the smallest effect the trial can find?
Data table
| Appointments per arm | Baseline 5% | Baseline 10% | Baseline 20% |
|---|---|---|---|
| 500 | 3.14 | 4.66 | 6.58 |
| 1000 | 2.37 | 3.44 | 4.76 |
| 1500 | 1.99 | 2.85 | 3.93 |
| 2000 | 1.75 | 2.50 | 3.42 |
| 2500 | 1.58 | 2.25 | 3.07 |
| 3000 | 1.46 | 2.06 | 2.81 |
| 3500 | 1.36 | 1.92 | 2.61 |
| 4000 | 1.28 | 1.80 | 2.45 |
| 4500 | 1.21 | 1.70 | 2.31 |
| 5000 | 1.15 | 1.62 | 2.19 |
| 5500 | 1.10 | 1.54 | 2.09 |
| 6000 | 1.06 | 1.48 | 2.01 |
| 6500 | 1.02 | 1.43 | 1.93 |
| 7000 | 0.98 | 1.38 | 1.86 |
| 7500 | 0.95 | 1.33 | 1.80 |
| 8000 | 0.92 | 1.29 | 1.74 |
| 8500 | 0.89 | 1.25 | 1.69 |
| 9000 | 0.87 | 1.22 | 1.64 |
| 9500 | 0.85 | 1.19 | 1.60 |
| 10000 | 0.83 | 1.16 | 1.56 |
| 10500 | 0.81 | 1.13 | 1.52 |
| 11000 | 0.79 | 1.10 | 1.49 |
| 11500 | 0.77 | 1.08 | 1.46 |
| 12000 | 0.76 | 1.06 | 1.43 |
| 12500 | 0.74 | 1.04 | 1.40 |
| 13000 | 0.73 | 1.02 | 1.37 |
| 13500 | 0.72 | 1.00 | 1.35 |
| 14000 | 0.70 | 0.98 | 1.32 |
| 14500 | 0.69 | 0.97 | 1.30 |
| 15000 | 0.68 | 0.95 | 1.28 |
| 15500 | 0.67 | 0.93 | 1.26 |
| 16000 | 0.66 | 0.92 | 1.24 |
| 16500 | 0.65 | 0.91 | 1.22 |
| 17000 | 0.64 | 0.89 | 1.20 |
| 17500 | 0.63 | 0.88 | 1.18 |
| 18000 | 0.62 | 0.87 | 1.17 |
| 18500 | 0.62 | 0.86 | 1.15 |
| 19000 | 0.61 | 0.85 | 1.14 |
| 19500 | 0.60 | 0.83 | 1.12 |
| 20000 | 0.59 | 0.82 | 1.11 |
Synthetic data. Source: projects/dna-reminder-experiment. Assumes independent appointments; multiply by the design effect if patients have several.
At a 10% baseline, 2,000 appointments per arm give 80% power to detect a fall of 2.5 points; 5,000 gets you to 1.62, 10,000 to 1.16, 20,000 to 0.82. If the detectable effect is bigger than anything the change could plausibly do, don’t run the trial as designed. Run it for longer, pool sites, or pick a more common outcome.
Write the analysis plan before the data exists
Decide these before the first patient is randomised, and date the document:
- The outcome and its exact definition. DNA per appointment, excluding appointments the hospital cancelled.
- The comparison and the threshold. Each text arm against no reminder, at 0.05 ÷ 3.
- The covariates you’ll adjust for, all measured before randomisation.
- A small number of subgroups, with the reason for each.
- When you’ll analyse. One look at the end, unless you’ve planned interim looks properly. Peeking and sequential testing explains why.
A plan written in advance is what stops you choosing, after the fact, the outcome definition or the subgroup that happens to look best.
Analyse by intention to treat
Intention to treat (ITT) compares everyone as randomised, whether or not they got what their arm intended. Some patients have no mobile number on record; some appointments are booked too late for a three-day text; some texts fail. ITT keeps them all, because a policy of “send reminders” would include them too. ITT answers the service’s question: what happens if we switch this on?
The tempting alternative is per-protocol: drop the people the intervention didn’t reach. The control arm keeps everyone, so you’re no longer comparing randomised groups. In the project, per-protocol made the one-text arm “significant” (p = 0.045) when the randomised comparison wasn’t (p = 0.096).
To ask what the change does for the people it reaches, stay inside the randomisation. Restrict every arm to a subgroup defined before randomisation, such as patients with a mobile number on record, or divide the ITT effect by the share of people the intervention reached (a complier-average estimate).
Report the effect and its interval
“Not significant” is not a finding. Report the effect, its 95% confidence interval, and the counts behind them.
From the project: two texts changed the DNA rate by −2.1 percentage points (95% CI −3.3 to −0.9). One text changed it by −1.0 points (95% CI −2.2 to +0.2), p = 0.096. The second result doesn’t say one text does nothing. It says the data fit anything from a 2.2-point fall, about as good as the best arm, to a small rise. The honest summary is “probably helps, not proven by this trial”.
Give absolute changes alongside relative ones. Going from 9.3% to 7.2% is a 23% relative fall, which sounds bigger than 2.1 points. Converting to counts helps managers most: about 210 fewer missed appointments per 10,000 booked.
Ethics and governance in a health setting
Randomising patients feels like a small operational tweak. The governance frameworks may not see it that way.
- It may count as research. In the UK, the Health Research Authority’s decision tool asks whether participants are randomised to different groups and whether care or services are allocated by randomisation. If both answers are yes, it classifies the study as research, and the next step is its separate tool on whether NHS Research Ethics Committee review is needed. The tool’s notes say clinical audit and service evaluation, by definition, don’t allocate care by randomisation. Talk to your research and development office early, before the allocation code is written.
- Information governance. The trial uses patient contact details and appointment data in a new way. Expect questions about the lawful basis, who sees the allocation, and how long the data is kept.
- Genuine uncertainty. Testing a new wording against the current one is easy to justify. Withdrawing a reminder that patients already get is harder. Compare against current practice unless there’s a reason not to.
- Who’s left out. If some patients can’t receive the intervention, such as those with no mobile number, say so, measure it, and plan for them.
- Wording that affects patients. A message about the cost of missed appointments is a behavioural nudge. Patient representatives should see it before patients do.
None of this is a reason not to experiment. It’s a reason to start the conversations before the start date.
A checklist
- The question, the outcome and one primary comparison, written down.
- The unit of randomisation, matched to how the change is delivered.
- The sample size from baseline, MDE, level and power, times the design effect.
- An allocation list generated with a fixed seed and saved before go-live.
- A dated analysis plan: ITT, covariates, subgroups, one analysis time.
- Governance classification agreed with the R&D office.
- A result reported as an effect with its interval, in counts a manager can use.
See it in a project
Tags
- experiments
- randomisation
- power
- sample-size
- statsmodels
- dna-rate