Prediction intervals for planners
Why a point forecast isn't a plan, how to build prediction intervals from backtest errors and quantile regression, how to check they hold, and how to choose and present a range for a staffing decision.
A forecast of 377 attendances next Monday is not something a rota manager can plan with. They need to know how bad a bad Monday gets. Staff to the point forecast every day and you will be short often: in the three-year backtest below, 48% of days came in above it. The number that matters for the plan is an upper percentile, and that comes from a prediction interval.
This article covers the difference between prediction and confidence intervals, two ways to build intervals that reflect how wrong your model has actually been, how to check that an 80% interval really contains the actual 80% of the time, and how to turn a range into a staffing decision that a non-analyst can follow.
The example
The data is synthetic, generated with a fixed seed by the script listed at the end: daily emergency department attendances from January 2021 to December 2025, around 300 a day. I planted a gentle upward trend, a weekly cycle with Monday the busiest day, a winter peak, a persistent day-to-day wobble that is twice as large in December to February, and occasional surge days that push attendances up by 10% to 25%.
The forecasting model is deliberately simple: a regression on trend, day of week and annual seasonality, plus a correction that carries the latest residual forward and lets it fade over the following days. The point here is the interval, not the model. The ED demand forecasting project compares serious models on similar data.
Prediction intervals and confidence intervals
These get confused all the time, and the difference is large.
- A confidence interval describes uncertainty about the average: where the regression line would sit if you could fit it on endless data. It shrinks as you add history.
- A prediction interval describes where a single future day will land. It includes the uncertainty about the average plus the day-to-day noise, and the noise doesn’t shrink however much history you have.
statsmodels returns both from get_prediction(...).summary_frame(alpha=0.2): mean_ci_lower and mean_ci_upper are the confidence interval, and obs_ci_lower and obs_ci_upper are the prediction interval. For the regression’s forecast of Monday 1 December 2025, the 80% confidence interval was 371 to 378. The 80% prediction interval was 339 to 410, more than ten times wider.
If someone hands you a tight band around a forecast, ask which one it is. A confidence interval drawn as a planning range will be wrong most days.
Why the textbook interval isn’t enough
That statsmodels prediction interval is only right if the model’s assumptions hold: errors that are independent, normally distributed and equally variable all year. Here, none of those is true. Errors persist from day to day, surges skew them upwards, and winter is noisier.
So rather than trust the formula, measure the errors. A rolling-origin backtest gives you exactly what you need: every week from January 2022 to December 2024, refit the model on the data before that week and forecast the next 28 days. That produced 4,284 forecasts, each with a horizon (1 to 28 days ahead) and an error (actual minus forecast), with a mean absolute error of 21.4 attendances.
Empirical intervals from backtest errors
The simplest honest interval takes the error quantiles for each horizon. For an 80% interval, take the 10th and 90th percentiles of the backtest errors at that horizon and add them to the new point forecast:
def empirical(backtest: pd.DataFrame, q: float) -> pd.Series:
"""Error quantile by horizon from the backtest."""
return backtest.groupby("h")["error"].quantile(q)
low = forecast["point"] + forecast["h"].map(empirical(backtest, 0.10))
high = forecast["point"] + forecast["h"].map(empirical(backtest, 0.90))
This gets three things right that the textbook interval misses. The interval is as wide as the model’s real errors, not its in-sample residuals. It can be lopsided when errors are skewed. And it can widen with horizon: here the 80% interval was 58.5 attendances wide one day ahead, 65.6 at seven days and 68.2 at 28, because the carried-forward correction helps most at short horizons.
What it can’t do is vary with anything other than horizon. If winter is harder to forecast, a winter day gets the same interval as a summer day.
Quantile regression
Quantile regression models a percentile directly instead of the mean. Fit it to the backtest errors with the things you expect to change the spread, here the log of the horizon and a winter flag, and you get error percentiles that depend on them:
def quantile_terms(frame: pd.DataFrame) -> pd.DataFrame:
return pd.DataFrame(
{"const": 1.0, "log_h": np.log(frame["h"]), "winter": frame["winter"]}, index=frame.index
)
x = quantile_terms(backtest)
models = {q: sm.QuantReg(backtest["error"], x).fit(q=q, max_iter=5000) for q in (0.1, 0.9)}
The fitted 90th-percentile model adds 25.6 attendances to the upper edge in winter. The 10th-percentile model moves the lower edge down by only 5.9. So in winter the interval gets wider and shifts upwards: the backtest shows the model tended to under-forecast in winter, and the quantile regression absorbs that as well as the extra noise. It would be better still to fix the bias in the point forecast, but the interval now tells the truth about it.
You can also fit quantile regression to the target itself, using the same features as the forecast, and skip the separate point model. Modelling the backtest errors has one advantage: you can wrap it around any point forecast you already have.
Checking coverage
An 80% interval should contain the actual value 80% of the time. That is testable, and it is the one check I would never skip. I made weekly forecasts through 2025 (1,316 forecast days, none of which the intervals had seen) and counted how often each interval contained the actual:
| 80% interval | All of 2025 | Winter | Rest of year |
|---|---|---|---|
| Textbook (statsmodels) | 81.3% | 66.2% | 84.3% |
| Backtest errors by horizon | 75.3% | 64.8% | 77.4% |
| Quantile regression | 77.4% | 80.6% | 76.7% |
Data table
| Method | Winter (Dec to Feb) | Rest of year |
|---|---|---|
| Textbook OLS interval | 66% | 84% |
| Backtest errors by horizon | 65% | 77% |
| Quantile regression | 81% | 77% |
Synthetic data. Source: projects/article-examples/stats/prediction_intervals.py.
The textbook interval looks perfect overall, at 81.3%. It isn’t. It is too wide for most of the year and too narrow in winter, and the two errors happen to cancel in the annual average. Winter is exactly when the rota is under most pressure, and that is when it contained the actual only two days in three.
The empirical interval fixes the width across horizons but not the seasonal problem. The quantile regression is the only one close to 80% in winter.
None of them hits 80% everywhere. 2025 came in below forecast on average (a mean error of −7.8 attendances), and a test year always has its own character. That is why coverage is something you monitor as the actuals arrive, not something you check once. Two habits help:
- Check coverage by segment, not just overall: season, day of week, horizon. An average can hide a failure where it matters most.
- Check each side separately when only one side matters. For staffing, the misses that hurt are days above the upper edge. In 2025, 4.9% of days came in above the quantile regression’s upper edge, against the 10% an 80% interval allows. For a rota, that is the safe direction.
Choosing 80% or 95% for a staffing decision
For staffing, the question isn’t “which interval?” but “which percentile do we plan to?” Two things make that easier.
An 80% interval’s upper edge is the 90th percentile. Ten per cent of days fall below the lower edge and 10% above the upper edge. A 95% interval’s upper edge is the 97.5th percentile.
The right percentile depends on costs. If one attendance with too few staff costs u (in delay, risk or agency spend) and one unit of spare capacity costs o, the percentile that minimises expected cost is u ÷ (u + o). Turn it around and each choice has a meaning:
| Plan to the | Implies a shortfall costs this many times a surplus |
|---|---|
| 75th percentile | 3 |
| 90th percentile (top of an 80% interval) | 9 |
| 97.5th percentile (top of a 95% interval) | 39 |
That is a conversation a service manager can have. For a core rota, where you can call in extra staff on the day, the 90th percentile is a defensible default. For something with no fallback, such as the minimum number of resus-trained nurses, the 97.5th percentile may be worth its price. In the December forecast, planning to the 97.5th percentile instead of the 90th meant cover for an average of 32.7 more attendances a day.
A forecast says how busy a day will be, not what a given plan does to waits; that is a what-if question, covered in simulation for what-if questions.
Presenting a range to non-analysts
Data table
| Date | Actual | Forecast | Forecast (low) | Forecast (high) |
|---|---|---|---|---|
| 2025-10-06 | 342 | |||
| 2025-10-07 | 306 | |||
| 2025-10-08 | 311 | |||
| 2025-10-09 | 310 | |||
| 2025-10-10 | 276 | |||
| 2025-10-11 | 296 | |||
| 2025-10-12 | 317 | |||
| 2025-10-13 | 333 | |||
| 2025-10-14 | 298 | |||
| 2025-10-15 | 297 | |||
| 2025-10-16 | 288 | |||
| 2025-10-17 | 366 | |||
| 2025-10-18 | 266 | |||
| 2025-10-19 | 300 | |||
| 2025-10-20 | 365 | |||
| 2025-10-21 | 302 | |||
| 2025-10-22 | 308 | |||
| 2025-10-23 | 303 | |||
| 2025-10-24 | 298 | |||
| 2025-10-25 | 262 | |||
| 2025-10-26 | 303 | |||
| 2025-10-27 | 341 | |||
| 2025-10-28 | 347 | |||
| 2025-10-29 | 298 | |||
| 2025-10-30 | 295 | |||
| 2025-10-31 | 342 | |||
| 2025-11-01 | 314 | |||
| 2025-11-02 | 321 | |||
| 2025-11-03 | 337 | |||
| 2025-11-04 | 300 | |||
| 2025-11-05 | 386 | |||
| 2025-11-06 | 318 | |||
| 2025-11-07 | 306 | |||
| 2025-11-08 | 307 | |||
| 2025-11-09 | 300 | |||
| 2025-11-10 | 343 | |||
| 2025-11-11 | 341 | |||
| 2025-11-12 | 318 | |||
| 2025-11-13 | 297 | |||
| 2025-11-14 | 344 | |||
| 2025-11-15 | 302 | |||
| 2025-11-16 | 304 | |||
| 2025-11-17 | 376 | |||
| 2025-11-18 | 357 | |||
| 2025-11-19 | 336 | |||
| 2025-11-20 | 364 | |||
| 2025-11-21 | 349 | |||
| 2025-11-22 | 339 | |||
| 2025-11-23 | 351 | |||
| 2025-11-24 | 396 | |||
| 2025-11-25 | 337 | |||
| 2025-11-26 | 332 | |||
| 2025-11-27 | 361 | |||
| 2025-11-28 | 358 | |||
| 2025-11-29 | 305 | |||
| 2025-11-30 | 288 | |||
| 2025-12-01 | 380 | 362 | 326 | 415 |
| 2025-12-02 | 351 | 342 | 305 | 395 |
| 2025-12-03 | 351 | 336 | 299 | 390 |
| 2025-12-04 | 348 | 336 | 299 | 390 |
| 2025-12-05 | 360 | 341 | 303 | 395 |
| 2025-12-06 | 320 | 321 | 283 | 375 |
| 2025-12-07 | 317 | 319 | 282 | 374 |
| 2025-12-08 | 442 | 377 | 340 | 432 |
| 2025-12-09 | 372 | 350 | 312 | 405 |
| 2025-12-10 | 367 | 341 | 303 | 396 |
| 2025-12-11 | 349 | 340 | 302 | 395 |
| 2025-12-12 | 324 | 344 | 306 | 399 |
| 2025-12-13 | 319 | 323 | 285 | 379 |
| 2025-12-14 | 337 | 322 | 284 | 377 |
| 2025-12-15 | 394 | 379 | 342 | 435 |
| 2025-12-16 | 341 | 352 | 314 | 408 |
| 2025-12-17 | 376 | 343 | 305 | 399 |
| 2025-12-18 | 353 | 342 | 304 | 397 |
| 2025-12-19 | 373 | 345 | 307 | 401 |
| 2025-12-20 | 339 | 325 | 287 | 381 |
| 2025-12-21 | 336 | 324 | 286 | 379 |
| 2025-12-22 | 408 | 381 | 343 | 437 |
| 2025-12-23 | 358 | 353 | 315 | 410 |
| 2025-12-24 | 321 | 344 | 306 | 400 |
| 2025-12-25 | 377 | 343 | 305 | 399 |
| 2025-12-26 | 412 | 347 | 308 | 403 |
| 2025-12-27 | 355 | 326 | 288 | 382 |
| 2025-12-28 | 306 | 325 | 286 | 381 |
Synthetic data. Source: projects/article-examples/stats/prediction_intervals.py.
For Monday 8 December 2025, the forecast was 377 attendances, with an 80% range of 340 to 432 and a 97.5th percentile of 466. Here is how I would say that to a rota manager:
We expect about 380 attendances next Monday. On eight Mondays in ten like this one, it will be between 340 and 430. If we staff for 430, we should be covered on nine Mondays in ten. Staffing for 470 covers all but about one Monday in 40.
A few habits make that work:
- Use frequencies, not percentages. “Nine Mondays in ten” lands better than “the 90th percentile”.
- Round. 377 suggests a precision the forecast doesn’t have.
- Show the band on every chart. A dashed line on its own invites people to treat it as a promise.
- Say what happened last time. That Monday came in at 442: above the 80% range and inside the 97.5th percentile. One day in ten is supposed to land above the range, and showing the misses alongside the hits is what builds trust in the next forecast.
Reproduce
The numbers come from projects/article-examples/stats/prediction_intervals.py, which simulates the data, runs the backtest and the 2025 test, fits the quantile regressions, and writes both charts:
uv run python projects/article-examples/stats/prediction_intervals.py
See it in a project
Tags
- prediction-intervals
- quantile-regression
- coverage
- backtesting
- ed-demand