Behnam Analytics

Writing Forecasting

Prediction intervals for planners

Why a point forecast isn't a plan, how to build prediction intervals from backtest errors and quantile regression, how to check they hold, and how to choose and present a range for a staffing decision.

Behnam Ebrahimi 8 min read

A forecast of 377 attendances next Monday is not something a rota manager can plan with. They need to know how bad a bad Monday gets. Staff to the point forecast every day and you will be short often: in the three-year backtest below, 48% of days came in above it. The number that matters for the plan is an upper percentile, and that comes from a prediction interval.

This article covers the difference between prediction and confidence intervals, two ways to build intervals that reflect how wrong your model has actually been, how to check that an 80% interval really contains the actual 80% of the time, and how to turn a range into a staffing decision that a non-analyst can follow.

The example

The data is synthetic, generated with a fixed seed by the script listed at the end: daily emergency department attendances from January 2021 to December 2025, around 300 a day. I planted a gentle upward trend, a weekly cycle with Monday the busiest day, a winter peak, a persistent day-to-day wobble that is twice as large in December to February, and occasional surge days that push attendances up by 10% to 25%.

The forecasting model is deliberately simple: a regression on trend, day of week and annual seasonality, plus a correction that carries the latest residual forward and lets it fade over the following days. The point here is the interval, not the model. The ED demand forecasting project compares serious models on similar data.

Prediction intervals and confidence intervals

These get confused all the time, and the difference is large.

  • A confidence interval describes uncertainty about the average: where the regression line would sit if you could fit it on endless data. It shrinks as you add history.
  • A prediction interval describes where a single future day will land. It includes the uncertainty about the average plus the day-to-day noise, and the noise doesn’t shrink however much history you have.

statsmodels returns both from get_prediction(...).summary_frame(alpha=0.2): mean_ci_lower and mean_ci_upper are the confidence interval, and obs_ci_lower and obs_ci_upper are the prediction interval. For the regression’s forecast of Monday 1 December 2025, the 80% confidence interval was 371 to 378. The 80% prediction interval was 339 to 410, more than ten times wider.

If someone hands you a tight band around a forecast, ask which one it is. A confidence interval drawn as a planning range will be wrong most days.

Why the textbook interval isn’t enough

That statsmodels prediction interval is only right if the model’s assumptions hold: errors that are independent, normally distributed and equally variable all year. Here, none of those is true. Errors persist from day to day, surges skew them upwards, and winter is noisier.

So rather than trust the formula, measure the errors. A rolling-origin backtest gives you exactly what you need: every week from January 2022 to December 2024, refit the model on the data before that week and forecast the next 28 days. That produced 4,284 forecasts, each with a horizon (1 to 28 days ahead) and an error (actual minus forecast), with a mean absolute error of 21.4 attendances.

Empirical intervals from backtest errors

The simplest honest interval takes the error quantiles for each horizon. For an 80% interval, take the 10th and 90th percentiles of the backtest errors at that horizon and add them to the new point forecast:

def empirical(backtest: pd.DataFrame, q: float) -> pd.Series:
    """Error quantile by horizon from the backtest."""
    return backtest.groupby("h")["error"].quantile(q)

low = forecast["point"] + forecast["h"].map(empirical(backtest, 0.10))
high = forecast["point"] + forecast["h"].map(empirical(backtest, 0.90))

This gets three things right that the textbook interval misses. The interval is as wide as the model’s real errors, not its in-sample residuals. It can be lopsided when errors are skewed. And it can widen with horizon: here the 80% interval was 58.5 attendances wide one day ahead, 65.6 at seven days and 68.2 at 28, because the carried-forward correction helps most at short horizons.

What it can’t do is vary with anything other than horizon. If winter is harder to forecast, a winter day gets the same interval as a summer day.

Quantile regression

Quantile regression models a percentile directly instead of the mean. Fit it to the backtest errors with the things you expect to change the spread, here the log of the horizon and a winter flag, and you get error percentiles that depend on them:

def quantile_terms(frame: pd.DataFrame) -> pd.DataFrame:
    return pd.DataFrame(
        {"const": 1.0, "log_h": np.log(frame["h"]), "winter": frame["winter"]}, index=frame.index
    )

x = quantile_terms(backtest)
models = {q: sm.QuantReg(backtest["error"], x).fit(q=q, max_iter=5000) for q in (0.1, 0.9)}

The fitted 90th-percentile model adds 25.6 attendances to the upper edge in winter. The 10th-percentile model moves the lower edge down by only 5.9. So in winter the interval gets wider and shifts upwards: the backtest shows the model tended to under-forecast in winter, and the quantile regression absorbs that as well as the extra noise. It would be better still to fix the bias in the point forecast, but the interval now tells the truth about it.

You can also fit quantile regression to the target itself, using the same features as the forecast, and skip the separate point model. Modelling the backtest errors has one advantage: you can wrap it around any point forecast you already have.

Checking coverage

An 80% interval should contain the actual value 80% of the time. That is testable, and it is the one check I would never skip. I made weekly forecasts through 2025 (1,316 forecast days, none of which the intervals had seen) and counted how often each interval contained the actual:

80% interval All of 2025 Winter Rest of year
Textbook (statsmodels) 81.3% 66.2% 84.3%
Backtest errors by horizon 75.3% 64.8% 77.4%
Quantile regression 77.4% 80.6% 76.7%
How often the 80% interval contained the actual, 2025 test forecastsA well-calibrated 80% interval should sit on the reference line
Data table
MethodWinter (Dec to Feb)Rest of year
Textbook OLS interval66%84%
Backtest errors by horizon65%77%
Quantile regression81%77%

Synthetic data. Source: projects/article-examples/stats/prediction_intervals.py.

The textbook interval looks perfect overall, at 81.3%. It isn’t. It is too wide for most of the year and too narrow in winter, and the two errors happen to cancel in the annual average. Winter is exactly when the rota is under most pressure, and that is when it contained the actual only two days in three.

The empirical interval fixes the width across horizons but not the seasonal problem. The quantile regression is the only one close to 80% in winter.

None of them hits 80% everywhere. 2025 came in below forecast on average (a mean error of −7.8 attendances), and a test year always has its own character. That is why coverage is something you monitor as the actuals arrive, not something you check once. Two habits help:

  • Check coverage by segment, not just overall: season, day of week, horizon. An average can hide a failure where it matters most.
  • Check each side separately when only one side matters. For staffing, the misses that hurt are days above the upper edge. In 2025, 4.9% of days came in above the quantile regression’s upper edge, against the 10% an 80% interval allows. For a rota, that is the safe direction.

Choosing 80% or 95% for a staffing decision

For staffing, the question isn’t “which interval?” but “which percentile do we plan to?” Two things make that easier.

An 80% interval’s upper edge is the 90th percentile. Ten per cent of days fall below the lower edge and 10% above the upper edge. A 95% interval’s upper edge is the 97.5th percentile.

The right percentile depends on costs. If one attendance with too few staff costs u (in delay, risk or agency spend) and one unit of spare capacity costs o, the percentile that minimises expected cost is u ÷ (u + o). Turn it around and each choice has a meaning:

Plan to the Implies a shortfall costs this many times a surplus
75th percentile 3
90th percentile (top of an 80% interval) 9
97.5th percentile (top of a 95% interval) 39

That is a conversation a service manager can have. For a core rota, where you can call in extra staff on the day, the 90th percentile is a defensible default. For something with no fallback, such as the minimum number of resus-trained nurses, the 97.5th percentile may be worth its price. In the December forecast, planning to the 97.5th percentile instead of the 90th meant cover for an average of 32.7 more attendances a day.

A forecast says how busy a day will be, not what a given plan does to waits; that is a what-if question, covered in simulation for what-if questions.

Presenting a range to non-analysts

Daily ED attendances with a 28-day forecast and its 80% intervalInterval from quantile regression on backtest errors
Data table
DateActualForecastForecast (low)Forecast (high)
2025-10-06342
2025-10-07306
2025-10-08311
2025-10-09310
2025-10-10276
2025-10-11296
2025-10-12317
2025-10-13333
2025-10-14298
2025-10-15297
2025-10-16288
2025-10-17366
2025-10-18266
2025-10-19300
2025-10-20365
2025-10-21302
2025-10-22308
2025-10-23303
2025-10-24298
2025-10-25262
2025-10-26303
2025-10-27341
2025-10-28347
2025-10-29298
2025-10-30295
2025-10-31342
2025-11-01314
2025-11-02321
2025-11-03337
2025-11-04300
2025-11-05386
2025-11-06318
2025-11-07306
2025-11-08307
2025-11-09300
2025-11-10343
2025-11-11341
2025-11-12318
2025-11-13297
2025-11-14344
2025-11-15302
2025-11-16304
2025-11-17376
2025-11-18357
2025-11-19336
2025-11-20364
2025-11-21349
2025-11-22339
2025-11-23351
2025-11-24396
2025-11-25337
2025-11-26332
2025-11-27361
2025-11-28358
2025-11-29305
2025-11-30288
2025-12-01380362326415
2025-12-02351342305395
2025-12-03351336299390
2025-12-04348336299390
2025-12-05360341303395
2025-12-06320321283375
2025-12-07317319282374
2025-12-08442377340432
2025-12-09372350312405
2025-12-10367341303396
2025-12-11349340302395
2025-12-12324344306399
2025-12-13319323285379
2025-12-14337322284377
2025-12-15394379342435
2025-12-16341352314408
2025-12-17376343305399
2025-12-18353342304397
2025-12-19373345307401
2025-12-20339325287381
2025-12-21336324286379
2025-12-22408381343437
2025-12-23358353315410
2025-12-24321344306400
2025-12-25377343305399
2025-12-26412347308403
2025-12-27355326288382
2025-12-28306325286381

Synthetic data. Source: projects/article-examples/stats/prediction_intervals.py.

For Monday 8 December 2025, the forecast was 377 attendances, with an 80% range of 340 to 432 and a 97.5th percentile of 466. Here is how I would say that to a rota manager:

We expect about 380 attendances next Monday. On eight Mondays in ten like this one, it will be between 340 and 430. If we staff for 430, we should be covered on nine Mondays in ten. Staffing for 470 covers all but about one Monday in 40.

A few habits make that work:

  • Use frequencies, not percentages. “Nine Mondays in ten” lands better than “the 90th percentile”.
  • Round. 377 suggests a precision the forecast doesn’t have.
  • Show the band on every chart. A dashed line on its own invites people to treat it as a promise.
  • Say what happened last time. That Monday came in at 442: above the 80% range and inside the 97.5th percentile. One day in ten is supposed to land above the range, and showing the misses alongside the hits is what builds trust in the next forecast.

Reproduce

The numbers come from projects/article-examples/stats/prediction_intervals.py, which simulates the data, runs the backtest and the 2025 test, fits the quantile regressions, and writes both charts:

uv run python projects/article-examples/stats/prediction_intervals.py