Writing Data analysis & statistics
Peeking and sequential testing
Why checking an experiment's p-value every week inflates false positives, a simulation that measures by how much, and the group sequential designs that let you look early without fooling yourself.
An experiment is running. Every Monday someone opens the dashboard and asks whether it’s working yet. If the rule is “stop as soon as p < 0.05”, the false-positive rate isn’t 5%. In the simulation below, a trial checked weekly for 12 weeks declared a difference that didn’t exist 20.7% of the time. Checked daily, 33.7%.
This article shows where that inflation comes from, how to avoid it with a fixed horizon, and how to look early properly when you need to, with Pocock, O’Brien-Fleming and alpha-spending boundaries.
One trial, checked every day
Here is one simulated A/A trial: two arms with exactly the same 10% DNA rate, 100 appointments per arm per clinic day, and a two-proportion z-test after every day.
Data table
| Clinic day | z statistic |
|---|---|
| 1 | 0.00 |
| 2 | -0.48 |
| 3 | 0.38 |
| 4 | 1.00 |
| 5 | 0.40 |
| 6 | 1.69 |
| 7 | 1.42 |
| 8 | 1.24 |
| 9 | 1.54 |
| 10 | 1.33 |
| 11 | 1.93 |
| 12 | 2.05 |
| 13 | 1.79 |
| 14 | 1.79 |
| 15 | 2.02 |
| 16 | 1.97 |
| 17 | 1.82 |
| 18 | 1.72 |
| 19 | 1.83 |
| 20 | 1.94 |
| 21 | 1.79 |
| 22 | 1.66 |
| 23 | 1.47 |
| 24 | 1.72 |
| 25 | 1.64 |
| 26 | 1.65 |
| 27 | 1.88 |
| 28 | 1.76 |
| 29 | 1.65 |
| 30 | 1.57 |
| 31 | 1.55 |
| 32 | 1.69 |
| 33 | 1.82 |
| 34 | 1.75 |
| 35 | 1.77 |
| 36 | 1.78 |
| 37 | 1.88 |
| 38 | 1.89 |
| 39 | 1.69 |
| 40 | 1.42 |
| 41 | 1.29 |
| 42 | 1.52 |
| 43 | 1.75 |
| 44 | 1.95 |
| 45 | 1.85 |
| 46 | 1.49 |
| 47 | 1.28 |
| 48 | 1.26 |
| 49 | 1.18 |
| 50 | 1.43 |
| 51 | 1.29 |
| 52 | 1.19 |
| 53 | 1.08 |
| 54 | 0.91 |
| 55 | 0.78 |
| 56 | 0.68 |
| 57 | 0.80 |
| 58 | 0.79 |
| 59 | 0.67 |
| 60 | 0.57 |
Synthetic data. Source: projects/dna-reminder-experiment. Trial 3 of 20,000: the first one in the run that crossed ±1.96 on some day.
The z statistic crossed ±1.96 on days 12, 15 and 16, and finished at 0.57 on day 60. Anyone checking daily would have stopped on day 12 and reported a difference between two identical arms. (I picked this trial as the first of the 20,000 in the run that crossed the line at some point. It isn’t unusual: a third of them did.)
The 5% in “p < 0.05” is a promise about one test, at a sample size fixed in advance. As data arrive, the z statistic wanders. Each look is another chance to catch it at an extreme, and stopping at the first crossing turns a passing fluctuation into a result. Consecutive looks share most of their data, so the inflation grows more slowly than it would for independent tests. But it keeps growing: a trial with no effect that keeps collecting data and keeps being checked is certain to cross 1.96 eventually.
How much peeking costs
The simulation runs 20,000 A/A trials of 60 clinic days (12 weeks of five days), 6,000 appointments per arm in total, both arms at a 10% DNA rate. For each number of equally spaced looks, from 1 to 60, it counts the trials that crossed |z| > 1.96 at any look.
The core is three small functions. Counts accumulate day by day, and the test at each look uses everything so far:
def pooled_z(a: np.ndarray, b: np.ndarray, n: np.ndarray) -> np.ndarray:
"""Two-proportion z statistic (pooled) from cumulative DNA counts `a`, `b` out of `n` each."""
pooled = (a + b) / (2 * n)
se = np.sqrt(pooled * (1 - pooled) * 2 / n)
return np.divide(b - a, n * se, out=np.zeros_like(se), where=se > 0)
def simulate_z_paths(
rng: np.random.Generator, trials: int, looks: int, per_look: int, p_a: float, p_b: float
) -> np.ndarray:
"""z statistic at each of `looks` equally spaced looks, one row per simulated trial."""
a = rng.binomial(per_look, p_a, (trials, looks)).cumsum(axis=1)
b = rng.binomial(per_look, p_b, (trials, looks)).cumsum(axis=1)
n = per_look * np.arange(1, looks + 1)
return pooled_z(a, b, n)
def crosses(z: np.ndarray, bounds: np.ndarray) -> np.ndarray:
"""First look at which |z| passes its boundary; -1 when it never does."""
over = np.abs(z) > bounds
return np.where(over.any(axis=1), over.argmax(axis=1), -1)
Data table
| Number of looks | 1.96 at every look | Pocock boundary | O'Brien-Fleming boundary |
|---|---|---|---|
| 1 | 5.2% | 5.2% | 5.2% |
| 2 | 8.6% | 5.1% | 5.1% |
| 3 | 10.7% | 4.9% | 5.1% |
| 4 | 12.7% | 4.9% | 5.1% |
| 5 | 14.1% | 5.0% | 5.2% |
| 6 | 15.7% | 5.0% | 5.1% |
| 10 | 19.3% | 5.0% | 4.9% |
| 12 | 20.7% | 5.0% | 5.1% |
| 15 | 22.5% | 4.9% | 5.2% |
| 20 | 24.7% | 5.0% | 5.0% |
| 30 | 28.2% | 4.9% | 5.0% |
| 60 | 33.7% | 4.9% | 5.0% |
Synthetic data. Source: projects/dna-reminder-experiment. Both arms at a 10% DNA rate, 100 appointments per arm per day for 60 days; looks equally spaced.
| Looks | Every | Trials declared “significant” |
|---|---|---|
| 1 | end only | 5.2% |
| 2 | 6 weeks | 8.6% |
| 5 | 12 days | 14.1% |
| 12 | week | 20.7% |
| 60 | day | 33.7% |
One look gives 5.2%, which is the promised 5% within simulation error. Two looks raise it to 8.6%. Weekly looks quadruple it.
The simplest fix: a fixed horizon
Decide the sample size before the trial starts, from a power calculation (A/B testing in operations shows how), and compare the arms once, at the end.
That doesn’t mean ignoring the trial while it runs. Watch the process: enrolment, balance between arms, texts delivered, data completeness. Those checks don’t test the hypothesis, so they cost nothing. What you don’t watch is the outcome comparison. A live dashboard with a running p-value is an invitation to stop early, so I wouldn’t build one.
The appointment reminder experiment shows the temptation. Its one-text arm crossed p < 0.05 in weeks 6 and 8 of a 16-week trial, then finished at p = 0.096. The effect was real (it was planted), but the trial was never likely to confirm it, and a weekly check would have announced it anyway.
Planned early looks: group sequential designs
Sometimes an early look is justified: an arm might be causing harm, a clear winner could be rolled out sooner, or the budget might not stretch. A group sequential design allows it. You fix the number of looks in advance and use stricter critical values at each one, so that the overall chance of a false positive stays at 5%.
The two classic boundaries, for five equally spaced looks:
- Pocock: the same critical value at every look, 2.41 instead of 1.96.
- O’Brien-Fleming: very strict early, relaxing towards the end. The critical value at look k of K is c × √(K / k), which for five looks gives 4.57, 3.23, 2.64, 2.29 and 2.04.
Data table
| Look | 1.96 at every look | Pocock | O'Brien-Fleming | O'Brien-Fleming-type spending |
|---|---|---|---|---|
| 1 | 1.96 | 2.41 | 4.57 | 4.38 |
| 2 | 1.96 | 2.41 | 3.23 | 3.11 |
| 3 | 1.96 | 2.41 | 2.64 | 2.55 |
| 4 | 1.96 | 2.41 | 2.29 | 2.25 |
| 5 | 1.96 | 2.41 | 2.04 | 2.07 |
Synthetic data. Source: projects/dna-reminder-experiment. Equally spaced looks; boundaries found by simulation.
I found the constants by simulation, taking the 95th percentile of the largest |z| across the looks in 200,000 simulated paths with no effect:
def pocock_constant(rng: np.random.Generator, looks: int) -> float:
"""The single critical value used at every look (Pocock), found by simulation."""
z = _brownian_z(rng, CALIBRATION_PATHS, looks)
return float(np.quantile(np.abs(z).max(axis=1), 1 - ALPHA))
They agree with the published values for five looks (2.41 for Pocock, 2.04 for the final O’Brien-Fleming boundary) to within simulation error. With the matching boundaries, the A/A trials’ false-positive rate stayed between 4.9% and 5.2% for every number of looks from 1 to 60: the flat lines in the chart above.
What each design costs
Early looks aren’t free. To compare the designs, I simulated trials with a real effect (10% falling to 8%), each with at most 3,205 appointments per arm, the size that gives a single-look trial 80% power, split into five looks of 641.
| Design | Power | Average share of the maximum sample used | Maximum sample needed for 80% power |
|---|---|---|---|
| Fixed, one look | 80.2% | 100% | 1.00 × |
| 1.96 at every look | 85.0% | 58% | not valid |
| Pocock | 71.2% | 70.5% | 1.22 × |
| O’Brien-Fleming | 78.8% | 80.0% | 1.03 × |
| O’Brien-Fleming-type spending | 78.7% | 78.6% | 1.04 × |
The naive design looks best on power, but only because its false-positive rate is 14.1%. Pocock stops earliest on average, but it needs a 22% bigger maximum sample to keep its power, and its final boundary of 2.41 creates an awkward case: a trial that ends at z = 2.2 has failed, though the same data would have passed a fixed design. O’Brien-Fleming barely spends anything early, costs about 3% in maximum sample, and its final boundary is close to 1.96, so the final analysis looks almost like a fixed-design one. For operational trials, O’Brien-Fleming with two to four looks is my default.
Alpha spending: looks you don’t have to schedule exactly
Pocock and O’Brien-Fleming boundaries assume the looks happen at the planned points. Operational trials rarely run to schedule: a clinic closes for a week, or the board meeting moves. Lan and DeMets (1983) solved this with the alpha-spending function. You fix how much of the 5% you’ll have spent by the time a fraction t of the planned information is in, and each boundary is set so that the chance of having crossed by then, with no effect, equals the amount spent.
The two common spending functions, as the R package gsDesign documents them:
- O’Brien-Fleming type: α(t) = 2 − 2Φ(Φ⁻¹(1 − α/2) / √t)
- Pocock type: α(t) = α ln(1 + (e − 1)t)
In my simulation, α is the two-sided 0.05. For five equal looks, the O’Brien-Fleming-type function then spends 0.00001, 0.0019, 0.0114, 0.0284 and finally 0.05, giving boundaries of 4.38, 3.11, 2.55, 2.25 and 2.07 in my simulation. Its power and cost are almost the same as the classic boundary (the last row of the table).
Packages don’t all read α that way. gsDesign treats α as one-sided, so a two-sided 5% design applies the function at 0.025 to each side. That spends even less early and gives boundaries of 4.88, 3.36, 2.68, 2.29 and 2.03 for the same five looks. Both versions keep the overall false-positive rate at 5%. Check which one your software uses before comparing boundaries.
The number and timing of looks no longer have to be fixed in advance, only the spending function and the maximum sample. The one rule that remains: the timing of a look can’t depend on what the results look like.
For a real trial I’d use a maintained package rather than my simulation. In R, gsDesign and rpact compute boundaries, spending functions and sample sizes exactly. Other principled approaches exist too, such as Bayesian monitoring with a decision rule fixed in advance. They share the same discipline: the stopping rule is written down before the data arrive.
Estimates from a trial that stopped early
A trial that stops at an interim look usually stops because its estimate is at a high point, so the effect it reports tends to be too big. In the reminder project, the week-8 estimates for two texts (−2.9 percentage points) and the cost message (−2.7) were both larger than the final estimates at week 16 (−2.1 and −1.8). Report the design and the look at which the trial stopped, and treat an early-stop estimate as an upper-leaning guess.
Rules I’d follow
- Fix the sample size and the analysis date before the start.
- Show process measures during the trial, not the outcome comparison.
- If interim looks are needed, write down their number, timing and boundaries first.
- Prefer O’Brien-Fleming-type boundaries: an early stop needs overwhelming evidence, and the final analysis is almost unchanged.
- Report the design alongside the result.
Reproduce
The simulation lives in trial_sequential.py in the reminder experiment’s code, and runs with the project:
uv run python projects/dna-reminder-experiment/run.py
See it in a project
Tags
- experiments
- sequential-testing
- false-positives
- obrien-fleming
- alpha-spending