Behnam Analytics

Work Healthcare BI

Appointment reminder experiment

A four-arm randomised trial of text reminders on synthetic outpatient appointments, sized from pre-trial data before any outcome existed, and analysed with patient-clustered intervals, intention to treat and cautious subgroups.

Synthetic data Every record here is generated. No real patient or organisational data is used.

DNA rate with two texts against none (95% CI −3.3 to −0.9)
−2.1 pts
Appointments per arm
4,956
Real coverage of '95%' intervals that ignore repeat appointments
92–93%
Text cost per DNA avoided with two texts
£2.50

Text reminders are one of the cheapest changes an outpatient service can make, and one of the easiest to get wrong in the analysis. This project runs a four-arm randomised experiment on synthetic outpatient appointments for a fictional acute hospital: no reminder, one text three days before, a second text the day before, and a text that names the cost of a missed appointment. It sizes the trial from pre-trial data before any outcome exists, then analyses it the way an analysis plan should.

All the data is synthetic, generated by the project code with a fixed seed. The effects are planted, so every estimate can be checked against the truth.

In short: two texts cut the DNA (did not attend) rate from 9.3% to 7.2%, a fall of 2.1 percentage points (95% CI 0.9 to 3.3). The cost message did about as well, at 1.8 points. One text probably helps, but this trial can’t show it: its interval runs from a 2.2-point fall to a 0.2-point rise. At 3p a text, reminders cost about £2.50 per DNA avoided, so money isn’t the hard part of the decision.

The question

A service manager wants to know three things: do texts reduce missed appointments, which wording works best, and is a second text worth sending? Whatever they choose will apply to every patient on the list, including those with no mobile number on record. That last point shapes the whole analysis.

Trial design

Arm What the patient gets
No reminder Nothing (the control)
One text A standard reminder three days before
Two texts The same reminder three days before, and another the day before
Cost message One text three days before that also says what a missed appointment costs

The cost-message arm borrows from two UK trials by Hallsworth and colleagues (PLOS ONE, 2015), where a reminder stating what a missed appointment costs the NHS had a DNA rate of 8.4%, against 11.1% for the existing message. I planted a smaller effect than that.

The analysis plan was fixed before any trial outcome existed:

  • Unit of randomisation: the patient. Someone with three appointments gets the same arm for all three, so nobody receives mixed messages. It also means their appointments aren’t independent, so every interval is clustered by patient.
  • Outcome: DNA, per appointment.
  • Primary analysis: intention to treat (ITT). Every appointment counts in its assigned arm, whether a text reached the patient or not.
  • Three comparisons, each text arm against no reminder, each tested at 0.05 ÷ 3 = 0.0167 (Bonferroni), so the chance of any false positive stays at or below 5%. I still show 95% intervals, because that’s what readers compare.
  • Adjustment covariates: age band, sex, deprivation (Index of Multiple Deprivation, IMD, quintile of the patient’s area), mobile number on record, and DNAs in the prior year. All recorded before randomisation.
  • No interim analysis. Sixteen weeks, one look at the end.

In a real service, randomising patients to different reminders counts as research under the Health Research Authority’s decision tool, so it would need the matching approvals; A/B testing in operations covers that side.

What’s in the data

Every pattern is a named constant at the top of trial_data.py.

Pattern What I planted What shows up
Baseline risk Higher for younger patients, more deprived areas, men, new appointments, long booking lead times, Mondays and Fridays Pre-trial DNA rate 13.5% at 18–29 and 7.8% at 65+; 11.9% in IMD quintile 1 and 8.1% in quintile 5
Patient tendency Each patient has their own propensity to miss, shared across their appointments (SD 0.9 on the log-odds scale) Repeat appointments are correlated: a design effect of 1.16
No mobile number 3% of 18–29s up to 22% of over-65s, scaled from 1.5 times in the most deprived fifth to 0.7 times in the least 10.3% of trial patients
Short notice About 5% of appointments booked under three days ahead, too late for the three-day text 4.5% of text-arm appointments
Failed delivery 3% of texts 2.5% of text-arm appointments
Reminder effects When a text arrives, on the log-odds scale: one text −0.25, two texts −0.33, cost message −0.38. Stronger for younger patients (1.4 times at 18–29, 0.5 times at 65+) and in deprived areas (1.2 times in quintile 1, 0.8 in quintile 5) Planted ITT effects of −1.6, −2.1 and −2.3 percentage points

Sizing the trial before looking

The sample size comes from 16 weeks of pre-trial appointments (19,357 of them, from 8,000 patients), with nothing from the trial itself. The baseline DNA rate was 10.0%. I set the smallest effect worth detecting at 2 percentage points: big enough to matter to a clinic, realistic for a reminder. Power 80%, two-sided test at 0.0167.

def appointments_per_arm(
    baseline: float, mde: float, alpha: float, power: float, design_effect: float = 1.0
) -> int:
    """Appointments needed in each arm to detect a fall of `mde` in a DNA rate of `baseline`."""
    effect = proportion_effectsize(baseline, baseline - mde)
    n = NormalIndPower().solve_power(effect_size=effect, alpha=alpha, power=power, ratio=1.0)
    return math.ceil(n * design_effect)

If appointments were independent, that’s 4,285 per arm. They aren’t: patients averaged 2.42 appointments in 16 weeks, and someone who misses one is more likely to miss the next. The pre-trial design effect (the variance of the DNA rate clustered by patient, divided by the binomial variance) was 1.16, so the plan needed 4,956 appointments per arm, about 8,194 patients in all.

Power by trial sizeBaseline DNA rate 10.0%, each arm against no reminder
Data table
Appointments per armFall of 1.5 pointsFall of 2 pointsFall of 3 points
2503%5%10%
5005%9%21%
7507%13%33%
10009%17%44%
125012%22%55%
150014%27%64%
175017%32%72%
200019%37%78%
225022%42%83%
250024%46%88%
275027%51%91%
300030%55%93%
325032%59%95%
350035%63%96%
375038%66%97%
400040%70%98%
425043%73%99%
450046%76%99%
475048%78%99%
500050%80%100%
525053%82%100%
550055%84%100%
575057%86%100%
600060%88%100%
625062%89%100%
650064%90%100%
675066%92%100%
700068%93%100%
725069%94%100%
750071%94%100%
775073%95%100%
800074%96%100%
825076%96%100%
850077%97%100%
875078%97%100%
900080%98%100%
925081%98%100%
950082%98%100%
975083%98%100%
1000084%99%100%
1025085%99%100%
1050086%99%100%
1075087%99%100%
1100088%99%100%
1125089%99%100%
1150090%99%100%
1175090%100%100%
1200091%100%100%

Synthetic data. Source: projects/dna-reminder-experiment. Two-sided test at 0.0167 (5% split across three comparisons), allowing for repeat appointments (design effect 1.16).

The planted one-text effect, 1.6 points, is below that 2-point target. Re-running the trial 2,000 times on the same patients, with fresh randomisation and fresh outcomes, it passed the 0.0167 threshold 63% of the time, against 89% for two texts and 94% for the cost message. An inconclusive result for one text was always a likely outcome.

Did randomisation work?

8,194 patients, 2,048 or 2,049 per arm, with 19,790 appointments (4,917 to 5,005 per arm).

At randomisation No reminder One text Two texts Cost message
Aged 18–29 17.5% 18.9% 18.2% 18.3%
Aged 65+ 27.7% 25.3% 25.3% 25.0%
IMD quintile 1 (most deprived) 23.4% 25.2% 25.0% 24.4%
No mobile number 10.7% 9.8% 10.2% 10.4%
Any DNA in the prior year 16.7% 15.5% 15.0% 18.0%
Appointments in the trial (mean) 2.41 2.40 2.44 2.41

Every standardised difference from the control arm is below 0.1; the largest is 0.06, for the share aged 65+ in the cost-message arm. I don’t test these differences for significance. Under randomisation, any imbalance is chance by definition; the useful question is whether it’s big enough to matter, and the standardised difference answers that. One gap does matter later: the cost-message arm had more patients with a prior DNA.

Results

DNA rate by trial armShare of appointments missed, with 95% intervals clustered by patient
Data table
ArmDNA rateDNA rate (low)DNA rate (high)
No reminder9.3%8.4%10.1%
One text8.3%7.4%9.1%
Two texts7.2%6.4%7.9%
Cost message7.5%6.7%8.3%

Synthetic data. Source: projects/dna-reminder-experiment. Intention to treat: every randomised appointment counts.

Arm DNA rate Change vs no reminder (95% CI) p Adjusted (95% CI) Planted
No reminder 9.3%
One text 8.3% −1.0 (−2.2 to +0.2) 0.096 −1.1 (−2.3 to +0.1) −1.6
Two texts 7.2% −2.1 (−3.3 to −0.9) 0.0004 −2.2 (−3.3 to −1.0) −2.1
Cost message 7.5% −1.8 (−2.9 to −0.6) 0.003 −2.0 (−3.1 to −0.8) −2.3

Changes are in percentage points, intention to treat, clustered by patient. Two texts and the cost message pass the pre-specified threshold. One text doesn’t.

The trial can’t rank the two stronger arms. The cost message minus two texts is +0.3 points (−0.8 to +1.4), and the planted difference is only 0.2 points the other way. Two texts minus one text is −1.1 points (−2.2 to +0.0, p = 0.059), against a planted −0.5.

Why the intervals are clustered

Treating the 19,790 appointments as independent gives intervals 6% to 7% narrower. In the 2,000 re-runs, the estimates’ real standard deviation was about 0.61 points; the naive standard errors averaged 0.56 to 0.57, the clustered ones 0.60 to 0.61. So the naive “95%” intervals contained the planted effect only 92% to 93% of the time, the clustered ones 94% to 95%.

What adjustment did

Adjusting for the baseline covariates narrowed the intervals by only about 2%: with a binary outcome and modest predictors, there’s little variance to remove. It did move the cost-message estimate from −1.8 to −2.0 points, because that arm happened to have more patients with a prior DNA (18.0% against 16.7%). Correcting chance imbalance in factors named before the trial is exactly what adjustment is for.

The logistic regression gives odds ratios, which managers can’t use. So the code predicts every appointment under each arm, averages, and subtracts, giving a risk difference; the interval comes from the delta method with the patient-clustered covariance:

def mean_and_gradient(arm: str) -> tuple[float, np.ndarray]:
    scenario = x.copy()
    scenario[:, columns] = 0.0
    if arm != CONTROL:
        scenario[:, fit.model.exog_names.index(ARM_TERMS[arm])] = 1.0
    p = 1.0 / (1.0 + np.exp(-scenario @ beta))
    return float(p.mean()), (p * (1 - p)) @ scenario / len(p)

Intention to treat, per-protocol, and the patients no text reaches

In the text arms, the three-day text reached 82.4% of appointments. The rest had no mobile number on record (10.6%), were booked too late (4.5%), or the text failed (2.5%).

One trial, six analysesEstimated change in DNA rate, percentage points, with 95% intervals
Data table
One text
AnalysisThis trialThis trial (low)This trial (high)Planted value
Appointments independent-1.0-2.10.1-1.6
Clustered by patient-1.0-2.20.2-1.6
Adjusted (logistic)-1.1-2.30.1-1.6
Mobile-number holders-1.2-2.50.0-1.8
Per-protocol-1.3-2.50.0-2.0
Receiving a text (IV)-1.2-2.70.2-2.0
Two texts
AnalysisThis trialThis trial (low)This trial (high)Planted value
Appointments independent-2.1-3.2-1.0-2.1
Clustered by patient-2.1-3.3-0.9-2.1
Adjusted (logistic)-2.2-3.3-1.0-2.1
Mobile-number holders-2.5-3.7-1.3-2.4
Per-protocol-2.6-3.8-1.5-2.4
Receiving a text (IV)-2.3-3.6-1.1-2.4
Cost message
AnalysisThis trialThis trial (low)This trial (high)Planted value
Appointments independent-1.8-2.9-0.7-2.3
Clustered by patient-1.8-2.9-0.6-2.3
Adjusted (logistic)-2.0-3.1-0.8-2.3
Mobile-number holders-2.0-3.2-0.7-2.6
Per-protocol-2.2-3.4-1.0-2.9
Receiving a text (IV)-2.2-3.6-0.7-2.9

Synthetic data. Source: projects/dna-reminder-experiment. Each analysis answers a slightly different question, so each has its own planted value.

ITT answers the service’s question: what happens to the DNA rate if we switch this on for everyone? Per-protocol drops the treated appointments no text reached, but keeps every control appointment, so it compares a selected group with an unselected one. For one text it gives −1.3 points with p = 0.045: a “significant” result that the randomised comparison doesn’t support.

There are two honest ways to ask what a text does for someone it reaches:

  • Restrict every arm to patients with a mobile number. That’s recorded before randomisation, so the comparison stays randomised: −1.2, −2.5 and −2.0 points, against planted values of −1.8, −2.4 and −2.6.
  • Divide the ITT effect by the share of appointments a text reached (the Wald instrumental-variable estimator): −1.2, −2.3 and −2.2 points, against planted effects among those reached of −2.0, −2.4 and −2.9.

In this synthetic world, per-protocol’s bias turned out small: across the re-runs it averaged −2.56 points for two texts against a true −2.42, and was within 0.01 points for the other arms. That’s because the patients texts miss have much the same baseline risk as those they reach (pre-trial DNA rate 10.4% without a mobile number, 10.0% with one): the older patients and short-notice bookings among them offset the extra risk. In a service where the unreachable patients are the likeliest to miss, the bias would be bigger, and nothing inside the per-protocol analysis would show it.

The no-mobile patients are also an equity question. They’re older and more often from deprived areas, and a text-only policy gives them nothing: their estimates in the subgroup chart scatter around zero, as planted. Chasing missing numbers, or a letter or call for those without one, belongs in the policy.

Subgroups, with caution

Effect on the DNA rate by subgroupText arm minus no reminder, percentage points, with 95% intervals
Data table
One text
SubgroupEstimateEstimate (low)Estimate (high)Planted effect
All patients-1.0-2.20.2-1.6
Aged 18-29-2.9-6.30.6-3.0
Aged 30-44-1.9-4.40.6-2.2
Aged 45-64-1.2-3.30.9-1.5
Aged 65+0.5-1.42.5-0.5
IMD 1 (most deprived)-0.6-3.42.3-2.1
IMD 2-0.6-3.11.9-1.8
IMD 3-3.1-5.7-0.5-1.5
IMD 40.0-2.72.6-1.3
IMD 5 (least deprived)-1.1-3.51.4-1.2
No mobile number0.8-3.35.00.0
Has a mobile number-1.2-2.50.0-1.8
Two texts
SubgroupEstimateEstimate (low)Estimate (high)Planted effect
All patients-2.1-3.3-0.9-2.1
Aged 18-29-3.4-6.80.0-4.0
Aged 30-44-3.0-5.4-0.5-2.9
Aged 45-64-3.3-5.2-1.4-1.9
Aged 65+0.2-1.72.2-0.7
IMD 1 (most deprived)-2.7-5.30.0-2.8
IMD 2-1.3-3.71.1-2.4
IMD 3-3.8-6.3-1.3-2.0
IMD 4-1.3-4.01.3-1.8
IMD 5 (least deprived)-1.2-3.81.4-1.6
No mobile number1.5-2.45.50.0
Has a mobile number-2.5-3.7-1.3-2.4
Cost message
SubgroupEstimateEstimate (low)Estimate (high)Planted effect
All patients-1.8-2.9-0.6-2.3
Aged 18-29-3.2-6.50.2-4.4
Aged 30-44-2.9-5.4-0.3-3.1
Aged 45-64-2.0-4.00.0-2.1
Aged 65+-0.2-2.11.7-0.8
IMD 1 (most deprived)-2.7-5.40.0-3.0
IMD 2-0.4-3.02.2-2.6
IMD 3-3.0-5.6-0.5-2.1
IMD 4-0.4-3.12.2-1.9
IMD 5 (least deprived)-2.1-4.40.2-1.7
No mobile number-0.1-3.93.60.0
Has a mobile number-2.0-3.2-0.7-2.6

Synthetic data. Source: projects/dna-reminder-experiment. Below zero means fewer missed appointments. With 12 rows per arm, expect about one interval in 20 to miss its planted value by chance alone.

The planted effects are larger for younger patients and in deprived areas. The age estimates have the right shape, from −2.9, −3.4 and −3.2 points at 18–29 to +0.5, +0.2 and −0.2 at 65+. But the joint test that the effects differ across age bands gives p = 0.36; across IMD quintiles, p = 0.86; with and without a mobile number, p = 0.30. An interaction test compares estimates that each rest on part of the data, so it needs a much bigger trial than the average effect does.

The deprivation pattern shows why subgroups need care. IMD quintile 3 had the largest estimated effects in every arm (−3.1, −3.8 and −3.0 points), though the planted gradient runs from quintile 1 to 5. None of the 36 subgroup intervals missed its planted value: the estimates are noisy, not wrong. Reporting “reminders work best in IMD quintile 3” would be reporting noise. I’d show every pre-specified subgroup, lead with the interaction test, and treat any pattern as a hypothesis for the next trial.

Why the early weeks would have misled

The estimated effect as the trial ranText arm minus no reminder, using every appointment up to that week, 95% band
Data table
One text
Trial weekEstimate so farEstimate so far (low)Estimate so far (high)
11.2-3.05.4
21.3-1.64.3
30.4-2.02.9
4-0.4-2.51.8
5-1.3-3.30.7
6-1.8-3.70.0
7-1.6-3.30.0
8-1.7-3.4-0.1
9-1.2-2.70.4
10-1.1-2.60.4
11-0.7-2.10.7
12-0.9-2.20.5
13-0.7-2.10.6
14-0.6-1.90.7
15-0.9-2.10.4
16-1.0-2.20.2
Two texts
Trial weekEstimate so farEstimate so far (low)Estimate so far (high)
1-0.8-4.73.0
20.5-2.43.3
3-0.5-2.91.8
4-0.9-3.01.2
5-2.4-4.3-0.5
6-2.4-4.1-0.6
7-2.4-4.0-0.8
8-2.9-4.5-1.3
9-2.8-4.3-1.3
10-2.8-4.2-1.4
11-2.2-3.6-0.9
12-2.1-3.4-0.8
13-2.0-3.2-0.7
14-1.8-3.0-0.6
15-2.0-3.1-0.8
16-2.1-3.3-0.9
Cost message
Trial weekEstimate so farEstimate so far (low)Estimate so far (high)
10.5-3.64.5
20.7-2.23.6
30.0-2.52.4
4-0.7-2.81.4
5-2.0-3.90.0
6-2.3-4.0-0.5
7-1.9-3.5-0.2
8-2.7-4.3-1.1
9-2.5-4.0-1.0
10-2.5-4.0-1.1
11-2.1-3.4-0.7
12-1.9-3.3-0.6
13-2.0-3.2-0.7
14-1.6-2.8-0.4
15-1.7-2.9-0.5
16-1.8-2.9-0.6

Synthetic data. Source: projects/dna-reminder-experiment. Intervals clustered by patient, not adjusted for repeated looks.

After weeks 1 and 2, the one-text arm had a higher DNA rate than the control (+1.2 and +1.3 points). By week 8, two texts and the cost message showed falls of 2.9 and 2.7 points, bigger than the final estimates and the planted effects. One text crossed p < 0.05 in weeks 6 and 8, then finished at p = 0.096. A manager checking every week at 5% would have declared it a success in week 6.

The fixed horizon is what lets the final p-values mean what they say. Peeking and sequential testing simulates how fast repeated looks inflate false positives, and covers designs that allow planned interim looks.

Is it worth it?

The assumptions are constants in trial_results.py, not findings: 3p per text segment, two segments for the longer cost message, and £50 of clinic time put to use for each DNA avoided. Set-up and admin costs are left out.

Per 10,000 booked appointments One text Two texts Cost message
Texts sent 8,495 17,516 8,431
Text cost £255 £525 £506
DNAs avoided (95% CI) 102 (−18 to 222) 210 (95 to 326) 177 (59 to 294)
Net value (95% CI) £4,837 (−£1,167 to £10,841) £9,987 (£4,220 to £15,755) £8,328 (£2,468 to £14,188)
Text cost per DNA avoided £2.50 £2.50 £2.86

Fewer texts than appointments go out because patients with no mobile number and short-notice bookings don’t get the three-day text. The second text costs £271 more per 10,000 appointments and avoids an estimated 108 more DNAs (−4 to 221).

At the point estimates, each arm pays for itself once an avoided DNA is worth more than £2.50 to £2.86. So the decision turns on what this sketch leaves out: set-up, opt-outs, how patients feel about a message that mentions cost, and whether freed slots are actually rebooked.

What I’d recommend

On this evidence, two texts as the default: it has the clearest effect, and the second text is close to free. The cost message did about as well with one text, and the trial can’t separate the two. The next trial should test the combination directly.

Limits

  • Planted, not observed. The effects, their gradients and who lacks a mobile number are my assumptions.
  • No rebooking. A DNA usually generates another appointment. I didn’t model that, or cancellations and opt-outs.
  • Per-appointment outcome. “Did this patient miss any appointment?” is a different question, with a different answer.
  • Sixteen weeks, one population, with no seasonality beyond week-to-week swings.
  • A cost-benefit sketch. The assumed £50 carries the whole net-value row.

What I’d do next

  • A 2 × 2 factorial trial: one or two texts, standard or cost wording. It answers both remaining questions with the same patients.
  • Blocked randomisation within age bands, to guarantee balance where the effect is expected to vary most.
  • A group sequential design with O’Brien-Fleming boundaries, if the service wants planned interim looks.
  • Model rebooking, and report a patient-level outcome alongside the appointment-level one.

Run it yourself

uv run python projects/dna-reminder-experiment/run.py

It sizes, simulates, analyses and re-runs the trial, rewrites every chart and download on this page, and prints the tables quoted here, in 20 to 50 seconds. A rerun gives byte-identical files. The full code is in the download below, with a README that explains each file.

Built with

  • Python
  • statsmodels
  • pandas
  • NumPy
  • SciPy

Downloads