Behnam Analytics

Work Healthcare BI

Emergency department demand forecasting

Daily attendance forecasts 28 days ahead from three models and an average, scored with a rolling-origin backtest across 2025 and given 80% intervals built from past errors.

Synthetic data Every record here is generated. No real patient or organisational data is used.

MASE of the best model over 28 days
0.79
Lower MAE than the seasonal naive baseline
28%
Actuals inside the 80% interval (target 80%)
76.4%
Noise-floor MAE; the best model reached 24.1
20.0

Rotas, escalation plans and bed plans for next month need a number for each day, and a range around it. This project forecasts daily emergency department attendances 28 days ahead for a fictional acute hospital, compares three models and a simple average, and scores them the way they would be used: twelve monthly forecasts across 2025, each made with only the data that existed on the day.

All the data is synthetic, generated by the project code with a fixed seed. No real hospital, patient or staff data is involved.

In short: an average of exponential smoothing and gradient boosting had the lowest error, a MASE of 0.79 and a mean absolute error 28% below a seasonal naive baseline. Its 80% intervals held on 76.4% of days, a little too narrow. And every model missed Christmas Day by 85 attendances or more.

The planning problem

A rota for the four weeks from 1 December has to be drafted in November. So the useful forecast is the one made at the end of a month for the month that follows, with nothing from that month in hand.

A planner needs three things from it:

  1. A daily number for each of the next 28 days, because staffing follows the weekly cycle and the bank holidays, not a monthly total.
  2. A range, so the plan can be set against a bad day rather than an average one. An 80% interval is a reasonable default: wide enough to be honest, narrow enough to act on.
  3. A warning when the forecast is weak. Christmas, bank holidays and heatwaves are exactly the days when a plan matters most, and they are where models do worst.

What’s in the data

Five years of daily attendances, 1 January 2021 to 31 December 2025, split into ambulance and walk-in arrivals. Every pattern below is planted on purpose and set as a named constant at the top of simulate.py, so the forecasts can be judged against known truth.

Pattern What I planted What shows up
Day of week Monday 13% above an average day, Saturday and Sunday 7% below Monday mean 342.0, Saturday 285.6, Sunday 285.3
Trend About 1.8% growth a year Yearly mean 291.3 in 2021, 321.1 in 2025
Winter pressure A 6% annual cycle peaking in mid-January, plus one respiratory wave each winter with a random peak (early December to late January) and a random size (2% to 8%) Winter peaks of different heights and timing
Bank holidays All 43 England and Wales bank holidays for 2021 to 2025, hard-coded from gov.uk. A bank holiday looks like a Sunday; the next working day gets an 8% catch-up; Christmas Day is a further 20% lower Includes the one-off 2022 Platinum Jubilee and State Funeral holidays and the 2023 Coronation
Heatwaves Eight spells of three to five days, three of them in summer 2025, peaking at +12% Short spikes each summer, weighted towards walk-ins
Data-feed outages Eight days recorded as missing, including one on a forecast origin (30 September 2025) and one inside a scored window (11 March 2025) Blank rows in the CSV
Noise A slow AR(1) disturbance on the log scale (φ = 0.85), then negative binomial counts with variance = mean + mean²/300 On the scored days of 2025, recorded counts sit 20.0 a day from the planted expected value on average

Ambulance arrivals are 28.1% of the total, a little higher in winter and lower during heatwaves. The CSV keeps the planted expected value for every day in a planted_expected column, which turns out to be the most useful column in the file: it shows how much of any miss was the model and how much was noise.

Emergency department arrivals, 2021 to 2025Average arrivals per day, by week
Data table
Week startingAll arrivalsBy ambulance
2021-01-04322101
2021-01-11329103
2021-01-18332102
2021-01-2530495
2021-02-0131294
2021-02-0830895
2021-02-1529890
2021-02-2230192
2021-03-0127881
2021-03-0827685
2021-03-1528585
2021-03-2227679
2021-03-2929285
2021-04-0527977
2021-04-1226768
2021-04-1928379
2021-04-2628177
2021-05-0327076
2021-05-1030581
2021-05-1728978
2021-05-2428872
2021-05-3126768
2021-06-0728771
2021-06-1428474
2021-06-2126464
2021-06-2829976
2021-07-0529472
2021-07-1227665
2021-07-1931875
2021-07-2629476
2021-08-0227874
2021-08-0927567
2021-08-1627769
2021-08-2326770
2021-08-3026866
2021-09-0628178
2021-09-1328076
2021-09-2029582
2021-09-2730380
2021-10-0430185
2021-10-1129484
2021-10-1828786
2021-10-2528882
2021-11-0128481
2021-11-0829987
2021-11-1529991
2021-11-2231897
2021-11-2930593
2021-12-0630086
2021-12-1329487
2021-12-2028989
2021-12-2729789
2022-01-03325105
2022-01-10356115
2022-01-17326102
2022-01-2430294
2022-01-31344101
2022-02-07338105
2022-02-1429095
2022-02-21319101
2022-02-2830794
2022-03-0730591
2022-03-1429791
2022-03-2129480
2022-03-2830488
2022-04-0429789
2022-04-1130792
2022-04-1829280
2022-04-2530481
2022-05-0226872
2022-05-0927273
2022-05-1628977
2022-05-2327768
2022-05-3029076
2022-06-0629474
2022-06-1328973
2022-06-2028377
2022-06-2728571
2022-07-0429571
2022-07-1129576
2022-07-1829574
2022-07-2527668
2022-08-0128976
2022-08-0829668
2022-08-1529574
2022-08-2228573
2022-08-2929878
2022-09-0530979
2022-09-1231385
2022-09-1930586
2022-09-2629783
2022-10-0331384
2022-10-1029783
2022-10-1734597
2022-10-2431288
2022-10-3132188
2022-11-0729185
2022-11-1430588
2022-11-2129788
2022-11-2828191
2022-12-05327103
2022-12-1232197
2022-12-19330104
2022-12-26338106
2023-01-02362119
2023-01-09358113
2023-01-1631399
2023-01-2333699
2023-01-30337105
2023-02-06323102
2023-02-1332496
2023-02-2032297
2023-02-2730295
2023-03-06345104
2023-03-1332192
2023-03-2031495
2023-03-2733090
2023-04-0329084
2023-04-1030086
2023-04-1731291
2023-04-2429582
2023-05-0129374
2023-05-0827774
2023-05-1527774
2023-05-2227976
2023-05-2926973
2023-06-0528664
2023-06-1231380
2023-06-1929474
2023-06-2628968
2023-07-0327966
2023-07-1030875
2023-07-1730778
2023-07-2429170
2023-07-3129575
2023-08-0727269
2023-08-1427074
2023-08-2128474
2023-08-2829370
2023-09-0431485
2023-09-1128677
2023-09-1829879
2023-09-2531781
2023-10-0231387
2023-10-0929382
2023-10-1632495
2023-10-2330688
2023-10-3031389
2023-11-0632388
2023-11-1332499
2023-11-2033097
2023-11-2733296
2023-12-0431595
2023-12-11332101
2023-12-18359112
2023-12-2531290
2024-01-0131293
2024-01-08337108
2024-01-15319102
2024-01-22333104
2024-01-2929988
2024-02-0532797
2024-02-1231998
2024-02-1932098
2024-02-26335105
2024-03-0432196
2024-03-1130685
2024-03-1831492
2024-03-2531490
2024-04-0131390
2024-04-0831087
2024-04-1531690
2024-04-2230282
2024-04-2931591
2024-05-0628775
2024-05-1329880
2024-05-2027770
2024-05-2729180
2024-06-0329369
2024-06-1028875
2024-06-1728672
2024-06-2427368
2024-07-0129070
2024-07-0828172
2024-07-1528767
2024-07-2227270
2024-07-2930475
2024-08-0527371
2024-08-1229671
2024-08-1930476
2024-08-2627670
2024-09-0230480
2024-09-0930086
2024-09-1630377
2024-09-2330979
2024-09-3030683
2024-10-0730186
2024-10-1429885
2024-10-2131994
2024-10-2830590
2024-11-0432797
2024-11-1132595
2024-11-18336103
2024-11-2532296
2024-12-0232295
2024-12-09336105
2024-12-16349114
2024-12-23340107
2024-12-30345108
2025-01-06328105
2025-01-13346111
2025-01-20326105
2025-01-27321100
2025-02-0330795
2025-02-1029589
2025-02-1732195
2025-02-2432195
2025-03-0333193
2025-03-1032094
2025-03-1732695
2025-03-2433494
2025-03-3133192
2025-04-0733494
2025-04-1433093
2025-04-2132093
2025-04-2830883
2025-05-0531285
2025-05-1232481
2025-05-1933088
2025-05-2633385
2025-06-0231586
2025-06-0930880
2025-06-1631273
2025-06-2329170
2025-06-3032180
2025-07-0730172
2025-07-1431576
2025-07-2128773
2025-07-2828176
2025-08-0429475
2025-08-1133481
2025-08-1828375
2025-08-2531082
2025-09-0129875
2025-09-0829379
2025-09-1529275
2025-09-2231688
2025-09-2932287
2025-10-0633493
2025-10-1334697
2025-10-2033998
2025-10-27344106
2025-11-03352104
2025-11-1034599
2025-11-1732794
2025-11-2434598
2025-12-01372109
2025-12-08346104
2025-12-15325102
2025-12-22309100

Synthetic data. Source: projects/ed-demand-forecasting. Weeks with an outage day average the days that were recorded.

Approach

Backtest design

Twenty-four forecast origins, one at each month end from 31 December 2023 to 30 November 2025. Each forecast covers the first 28 days of the following month. The twelve origins that forecast 2024 are for calibrating the prediction intervals and for design choices. The twelve that forecast 2025 are the ones scored: 335 days per model, because the outage on 11 March has no actual to score against.

Every model receives the history up to its origin and nothing else. Outage days in that history are filled with the same weekday one week earlier, which only looks backwards. The rolling-origin backtesting article explains why this beats a single train/test split.

Models

  1. Seasonal naive. Repeat the last observed week. It is the baseline any other model has to beat.
  2. Exponential smoothing (ETS). Holt-Winters from statsmodels, with a damped additive trend and multiplicative weekly seasonality, refitted at every origin. It knows nothing about bank holidays or the time of year.
  3. Gradient boosting (GBM). scikit-learn’s HistGradientBoostingRegressor as a direct multi-horizon model: one regressor, with the horizon as a feature. Every (origin, horizon) pair inside the history becomes a training row, 29,526 of them at the first 2025 origin and 38,878 at the last.
  4. Average of ETS and GBM. The mean of the two point forecasts.

The GBM predicts the log of attendances relative to the 28-day mean at the origin, so the trees never have to extrapolate the trend. Its features are the horizon, the day of week, flags for bank holidays, the working day after one, and Christmas Day, and four ratios to that 28-day level. Every lagged value is chosen so that it existed at the origin:

def direct_features(y, calendar, origin, horizon):
    cumsum = np.concatenate(([0.0], np.cumsum(y)))
    target = origin + horizon
    level = _window_mean(cumsum, origin, LEVEL_DAYS)
    # The four most recent observed days with the same weekday as the target.
    back = target - WEEK * np.ceil(horizon / WEEK).astype(int)
    same_weekday = (y[back] + y[back - 7] + y[back - 14] + y[back - 21]) / 4
    last_week = _window_mean(cumsum, origin, WEEK)
    last_year_week = _window_mean(cumsum, target - YEAR + 3, WEEK)  # centred on target - 364
    last_year_month = _window_mean(cumsum, target - YEAR + 14, LEVEL_DAYS)
    ...

For a target ten days ahead, “same weekday” means 14 days before the target, not 7, because 7 days before the target is still in the future. At forecast time y stops at the origin, so a feature that reached past it would raise an IndexError instead of leaking quietly.

An earlier version also had day of year as a feature. It did worse on the 2024 origins: with three or four winters of history, the trees learnt the timing of past respiratory waves, and that timing moves every year. The feature set above is the one that scored best on the 2024 origins. To be straight about it: I had also seen the earlier version’s 2025 scores before making that change, so the 2025 results aren’t a perfectly clean hold-out for the gradient boosting model. The choice holds up on 2024 alone, which is why I kept it.

Prediction intervals

The intervals are empirical, not model-based. For each origin, I take the log residuals, log(actual ÷ forecast), from the previous 12 origins, split them by model and horizon week, and use the 10th and 90th percentiles as multiplicative bounds. The first 2025 interval therefore uses only 2024 errors, and no interval ever sees an actual from its own forecast window.

past = out[out["origin"].isin(origins[i - INTERVAL_WINDOW : i]) & (out["date"] <= origin)]
quantiles = past.groupby(["model", "week"])["log_residual"].quantile(list(INTERVAL))

Prediction intervals for planners covers this method and quantile regression in more depth.

Results

Model MAE RMSE MAPE MASE Bias 80% coverage Mean width
Seasonal naive 33.5 44.1 10.7% 1.09 +2.0 74.3% 100.7
Exponential smoothing 24.9 32.6 7.9% 0.81 +1.1 75.5% 76.4
Gradient boosting 24.2 31.2 7.6% 0.79 −4.0 75.8% 70.5
Average of ETS and GBM 24.1 31.4 7.6% 0.79 −1.4 76.4% 72.3

Errors are in attendances per day. Bias is the mean of forecast minus actual, so a negative number means under-forecasting. MASE is scaled by the in-sample error of a one-week seasonal naive at each origin; the forecast accuracy metrics article explains that choice.

Mean absolute scaled error by model12 monthly origins in 2025, days 1 to 28. Lower is better.
Data table
ModelMASE
Seasonal naive1.09
Exponential smoothing0.81
Gradient boosting0.79
Average of ETS and GBM0.79

Synthetic data. Source: projects/ed-demand-forecasting.

The seasonal naive scores above 1 because its denominator is a one-week-ahead naive forecast, and forecasting up to four weeks ahead is harder. The other three beat it by a wide margin. The average edges the GBM, 0.787 against 0.791, but that gap is far smaller than the month-to-month swings, and the GBM won six of the twelve months on its own. I’d call them tied.

Accuracy by horizon

Error by forecast horizonMean absolute error by days ahead, 12 origins in 2025
Data table
Days aheadSeasonal naiveExponential smoothingGradient boostingAverage of ETS and GBM
129.225.419.822.6
247.531.630.130.8
327.319.227.823.5
437.124.024.324.1
530.728.632.128.2
618.915.316.715.9
744.331.929.930.8
823.827.622.525.0
937.516.915.416.2
1020.815.813.913.0
1137.927.126.326.1
1227.723.325.924.4
1336.432.330.830.5
1439.024.623.323.5
1529.325.625.625.5
1641.721.224.823.0
1735.628.626.427.5
1833.722.616.119.3
1935.325.829.327.1
2030.626.024.325.1
2138.320.819.018.7
2232.925.625.624.9
2332.018.917.217.9
2438.126.924.425.7
2542.432.332.731.0
2628.930.026.126.6
2730.227.925.526.3
2831.822.422.522.3

Synthetic data. Source: projects/ed-demand-forecasting. Each point averages at most 12 days, so expect wobble.

Error barely grows across the 28 days. The average’s MAE by horizon week runs 25.1, 22.6, 23.7 and 24.9. The level of demand moves slowly, and most of the remaining error is day-to-day noise that no model can predict at any horizon.

That noise has a size. A forecast that knew the planted expected value for every day, with no error at all, would still have an MAE of 20.0 against the recorded counts. The best model reaches 24.1. Measured against the planted expected value instead of the noisy actual, the GBM’s own error is 13.0 and the average’s 13.1, against 14.7 for ETS and 25.7 for the seasonal naive.

Intervals

How often the 80% interval contained the actualShare of scored days in 2025 inside the interval
Data table
ModelCoverage
Seasonal naive74.3%
Exponential smoothing75.5%
Gradient boosting75.8%
Average of ETS and GBM76.4%

Synthetic data. Source: projects/ed-demand-forecasting. Intervals come from past backtest residuals by horizon week.

Every model’s 80% interval held on 74% to 76% of days, so the intervals are a few points too narrow. The 2024 residuals that calibrated them came from a calmer year: on the 2024 origins the average’s MAE was 21.9, against 24.1 in 2025, and 2025 had two months where every model was wrong in the same direction (below). For a planner the practical reading is that the interval missed on 23.6% of days rather than the 20% it promised. Widening it, or refreshing it after a bad month, would be the fix.

Where the best models still fail

Because the data is synthetic, each miss can be split into the model’s own error (forecast minus the planted expected value) and noise (actual minus expected). The chart shows only the model’s part.

Model error on ordinary and special daysMean absolute gap between forecast and the planted expected value, 2025
Data table
Day type (days)Seasonal naiveExponential smoothingGradient boostingAverage of ETS and GBM
Ordinary days (313)24.713.712.412.3
Bank holidays (8)59.346.422.234.0
Working day after a bank holiday (5)26.421.512.216.5
Planted heatwave days (9)28.414.927.120.9

Synthetic data. Source: projects/ed-demand-forecasting. Heatwaves are planted in the data but unknown to every model.

Christmas Day. Actual 238, planted expected 241.9. The forecasts were 323 (GBM), 329 (average), 335 (ETS) and 341 (seasonal naive). The GBM has a Christmas flag, but it can’t use it: min_samples_leaf is 200, and the training window holds three Christmas Days, 84 rows across the 28 horizons. The trees can’t give Christmas a leaf of its own.

Forecast for December 2025, made on 30 NovemberAverage of ETS and GBM, dashed, with its 80% interval shaded
Data table
DateActualForecast (Average of ETS and GBM)Forecast (Average of ETS and GBM) (low)Forecast (Average of ETS and GBM) (high)
2025-11-10407
2025-11-11400
2025-11-12328
2025-11-13308
2025-11-14362
2025-11-15313
2025-11-16296
2025-11-17362
2025-11-18372
2025-11-19338
2025-11-20311
2025-11-21293
2025-11-22304
2025-11-23310
2025-11-24417
2025-11-25319
2025-11-26349
2025-11-27341
2025-11-28372
2025-11-29300
2025-11-30317
2025-12-01449386333432
2025-12-02379353304394
2025-12-03317338292378
2025-12-04375341294381
2025-12-05403338291377
2025-12-06338319275357
2025-12-07343317274355
2025-12-08411380344431
2025-12-09360354320401
2025-12-10325338305383
2025-12-11345335303380
2025-12-12352336304381
2025-12-13301318287360
2025-12-14326318288361
2025-12-15438376334424
2025-12-16330356316401
2025-12-17317338300381
2025-12-18326334297377
2025-12-19269335298378
2025-12-20271318282358
2025-12-21325320284361
2025-12-22355388347440
2025-12-23343354317401
2025-12-24317338303383
2025-12-25238329295373
2025-12-26298329295373
2025-12-27284319286362
2025-12-28325317284359

Synthetic data. Source: projects/ed-demand-forecasting. Gaps in the actual line are data-feed outage days.

Bank holidays. Here the flags earn their place. The GBM’s own error on the eight 2025 bank holidays was 22.2, against 46.4 for ETS, which forecasts a bank holiday as if it were an ordinary weekday. Averaging the two dilutes the gain to 34.0, so the “best” model is worse than the GBM alone on exactly the days a planner worries about.

Heatwaves. No model knows about heat. On 12 to 14 August the actuals were 382, 368 and 362, against the average’s 305, 290 and 288. ETS’s smaller heatwave error in the chart is luck, not insight: its forecasts sat higher than the GBM’s on all nine heatwave days, for reasons that have nothing to do with heat.

Forecast for August 2025, made on 31 JulyAverage of ETS and GBM, dashed, with its 80% interval shaded
Data table
DateActualForecast (Average of ETS and GBM)Forecast (Average of ETS and GBM) (low)Forecast (Average of ETS and GBM) (high)
2025-07-11323
2025-07-12284
2025-07-13334
2025-07-14358
2025-07-15302
2025-07-16321
2025-07-17315
2025-07-18289
2025-07-19324
2025-07-20296
2025-07-21323
2025-07-22305
2025-07-23282
2025-07-24285
2025-07-25314
2025-07-26255
2025-07-27248
2025-07-28267
2025-07-29307
2025-07-30290
2025-07-31286
2025-08-01284287247321
2025-08-02264270232302
2025-08-03272268231300
2025-08-04354325280364
2025-08-05320303261339
2025-08-06300290250325
2025-08-07289289249324
2025-08-08275288258314
2025-08-09292272244297
2025-08-10227270242295
2025-08-11380326292356
2025-08-12382305273333
2025-08-13368290260316
2025-08-14362288258314
2025-08-15263288256325
2025-08-16269272242308
2025-08-17312270240305
2025-08-18292317282358
2025-08-19274306272345
2025-08-20281291258329
2025-08-21299291258328
2025-08-22301288260317
2025-08-23266276250304
2025-08-24268275248303
2025-08-25335316286348
2025-08-26432329298362
2025-08-27271293265323
2025-08-28265292264322

Synthetic data. Source: projects/ed-demand-forecasting. Gaps in the actual line are data-feed outage days.

Noise that looks like a miss. On 26 August, the day after the summer bank holiday, 432 people attended. The planted expected value was 349, and the GBM forecast 351. With real data this would look like an 81-attendance model failure, and someone would be tempted to fix the model. It was noise.

Level shifts. February was over-forecast by 23.5 to 27.0 a day by every model. The planted winter wave faded after January, the planted expected level fell from 334.7 to 314.7, and the models carried January’s level forward. October was under-forecast by 27.4 to 30.3 a day by ETS, the GBM and the average, because the planted expected level jumped from 304.9 in September to 334.1. The seasonal naive, which just copies the last week of September, was least wrong in October, the only month it won.

Limits

  • The seasonal shape is planted, not observed. I put the winter peak into attendances. NHS England’s monthly A&E report for June 2016 notes that A&E attendances usually peak in the summer months, while emergency admissions generally peak in winter. For a real hospital, the seasonal shape has to come from its own data.
  • No COVID-era disruption. Real 2021 data would need special handling, or excluding from training.
  • Twelve origins is a small test. 335 scored days per model is enough to separate the seasonal naive from the rest, not enough to separate the GBM from the average.
  • Thin interval calibration. Each horizon week’s interval rests on about 84 past residuals, 12 origins by 7 days.
  • One daily series. Rotas need arrivals by hour, and bed plans need admissions, not attendances.

What I’d do next

  • Use each model where it’s strong. Take the GBM on flagged days and the average elsewhere, and test that rule on the 2024 origins before believing it.
  • Handle Christmas explicitly. A multiplicative adjustment estimated from past Christmases, or a smaller minimum leaf size, checked on the 2024 origins so the rest of the year doesn’t pay for it.
  • Add a weather forecast. Temperature for heatwaves, using forecast temperatures available at the origin, not observed ones, to avoid leaking the future.
  • Adaptive intervals. Widen the band after a month where coverage slips, or fit quantile GBMs with loss="quantile" and compare their pinball loss with the empirical intervals.
  • More origins. Weekly origins would give about four times as many forecasts at each horizon, and a firmer ranking.
  • Turn it into a bed plan. Forecast admissions instead of attendances and feed them to the bed occupancy simulation as its arrival rate.
  • Keep the parts coherent. Separate forecasts for ambulance and walk-in arrivals won’t add up to the total; hierarchical forecast reconciliation makes them agree.

Run it yourself

uv run python projects/ed-demand-forecasting/run.py

It regenerates the data, runs all 24 origins, rewrites every chart and download on this page, and prints the tables quoted here, in about 20 to 30 seconds. A rerun gives byte-identical files. The full code is in the download below, with a README that explains each file.

Built with

  • Python
  • pandas
  • statsmodels
  • scikit-learn
  • NumPy

Downloads