Backtesting forecasts with rolling origins
Why one train/test split misleads, how rolling-origin evaluation works, expanding versus sliding windows, horizon-aware features, a compact Python loop, and the mistakes that make backtests flatter a model.
A forecast should be judged on one question: how wrong was it when it was made with only the data available at the time? A single train/test split answers that question once. A rolling-origin backtest answers it many times, at different points in the year, and that difference decides which model you end up trusting.
The examples here come from my emergency department demand forecasting project: synthetic daily attendances, four models, and twelve monthly forecasts across 2025, each 28 days long.
Why one split misleads
The usual first attempt is to hold out the last month, fit on everything before it, and compare models on that month. Suppose the month you held out was October 2025. On that month alone, the seasonal naive (repeat last week) had the lowest mean absolute scaled error (MASE) of the four models:
| Model | MASE, October only | MASE, all 12 months | Months won |
|---|---|---|---|
| Seasonal naive | 1.05 | 1.09 | 1 |
| Exponential smoothing | 1.18 | 0.81 | 4 |
| Gradient boosting | 1.07 | 0.79 | 6 |
| Average of ETS and GBM | 1.12 | 0.79 | 1 |
A single October split would have told you to ship the baseline. Across twelve origins, the baseline is the worst model by a distance, and it won only that one month.
Data table
| Forecast month, 2025 | Seasonal naive | Exponential smoothing | Gradient boosting | Average of ETS and GBM |
|---|---|---|---|---|
| Jan | 1.08 | 0.58 | 0.55 | 0.52 |
| Feb | 1.20 | 1.09 | 1.06 | 1.07 |
| Mar | 0.98 | 0.61 | 0.69 | 0.64 |
| Apr | 0.76 | 0.63 | 0.58 | 0.59 |
| May | 0.80 | 0.64 | 0.77 | 0.64 |
| Jun | 1.37 | 0.98 | 0.77 | 0.83 |
| Jul | 0.88 | 0.76 | 0.81 | 0.78 |
| Aug | 1.29 | 0.91 | 0.90 | 0.90 |
| Sep | 1.36 | 0.64 | 0.65 | 0.65 |
| Oct | 1.05 | 1.18 | 1.07 | 1.12 |
| Nov | 1.33 | 0.75 | 0.73 | 0.74 |
| Dec | 1.03 | 0.97 | 0.92 | 0.94 |
Synthetic data. Source: projects/ed-demand-forecasting.
The chart shows why. Each month has its own conditions: a respiratory wave fading in February, a jump in level in October, Christmas in December. A single split samples one of those conditions and calls it the truth. The gradient boosting model’s monthly MASE ranged from 0.55 to 1.07. Which of those you report depends entirely on which month you happened to hold out.
A single split has two further problems. It gives you one error per horizon, so you can’t see how error grows with lead time. And it gives you no spread of errors, which is exactly what you need to build a prediction interval.
How rolling-origin evaluation works
Pick a series of forecast origins. At each one, fit the model on data up to the origin, forecast the horizon you care about, and store every forecast with its origin, target date and horizon. Then move the origin forward and repeat. Hyndman and Athanasopoulos call this evaluation on a rolling forecasting origin, because the origin rolls forward in time.
origin 1 [train ===============]|--- 28-day forecast ---|
origin 2 [train ====================]|--- 28-day forecast ---|
origin 3 [train =========================]|--- 28-day forecast ---|
time →
Three choices shape the backtest, and each should mirror how the forecast will be used:
- Origin spacing. If planners get a forecast at each month end, put origins at month ends. More frequent origins give more test days, but the test should still look like the real process.
- Horizon. Match the decision’s lead time. A rota drafted four weeks ahead needs 28-day forecasts, and it needs them scored at 28 days, not at one.
- How many origins. Enough to cover the seasons at least once. Twelve monthly origins cover a year, with 335 scored days per model in my project once an outage day is dropped.
The output is a long table of forecasts. Every metric, by model, by horizon, by month or by kind of day, is a group-by on that table.
Expanding or sliding windows
At each origin you choose how much history to train on.
An expanding window uses everything from the start of the series up to the origin. Estimates are more stable, and a model with annual features gets more years to learn from. The assumption is that old data is still relevant. My project uses expanding windows for both the exponential smoothing and the gradient boosting model.
A sliding window keeps a fixed length of the most recent history, say two years. It adapts faster when the process changes, such as a new urgent treatment centre diverting minor injuries or a change in how attendances are counted. The price is less data at every fit.
If you use scikit-learn, TimeSeriesSplit gives you expanding windows by default; set max_train_size for a sliding one, test_size for the horizon and gap to leave out samples between training and test. It splits by position, so for daily data with month-end origins I find a small explicit loop clearer. That is the next section.
When in doubt, run both. If a sliding window wins, something in the older data no longer describes the present, and that’s worth knowing on its own.
A compact implementation
This is the whole loop. It works with any function that takes a history and returns horizon forecasts:
import numpy as np
import pandas as pd
def rolling_origin(y, origins, horizon, fit_predict, window=None):
"""Forecast `horizon` days from each origin, using only data up to that origin."""
frames = []
for origin in origins:
history = y.loc[:origin]
if window is not None: # sliding window: keep only the most recent days
history = history.iloc[-window:]
history = history.fillna(history.shift(7)) # fill gaps from the past only
dates = pd.date_range(origin + pd.Timedelta(days=1), periods=horizon, freq="D")
frames.append(
pd.DataFrame(
{
"origin": origin,
"date": dates,
"horizon": np.arange(1, horizon + 1),
"forecast": fit_predict(history, horizon),
"actual": y.reindex(dates).to_numpy(),
}
)
)
return pd.concat(frames, ignore_index=True).dropna(subset=["actual"])
def seasonal_naive(history, horizon, period=7):
return np.resize(history.to_numpy()[-period:], horizon)
y = pd.read_csv("ed-attendances-daily.csv", index_col="date", parse_dates=True)["attendances"]
origins = pd.date_range("2024-12-31", "2025-11-30", freq="ME") # month ends
bt = rolling_origin(y, origins, horizon=28, fit_predict=seasonal_naive)
bt["abs_error"] = (bt["forecast"] - bt["actual"]).abs()
Run on the project’s daily attendances file, this gives the seasonal naive an MAE of 33.5 over 335 days, the same as the full project. Its monthly MAE ranged from 23.1 to 41.9, which is the single-split problem again in one line of output.
Two details matter more than they look. Missing days in the history are filled from the same weekday a week earlier, which only looks backwards. And missing actuals are dropped, not filled, because a filled value is your own guess, and scoring a forecast against a guess tells you nothing.
Horizon-aware features
Rolling origins fix the evaluation. They don’t fix the features, and features are where most leakage hides.
Say you build a lag-1 feature with df["lag_1"] = df["y"].shift(1) on the full series, then train a model and forecast 28 days ahead. On day 10 of the forecast, lag 1 is day 9, which you don’t know at the origin. In the backtest the column is quietly filled with the real value, and your 28-day accuracy looks like 1-day accuracy.
There are three honest ways out:
- Recursive. Forecast one day, feed the forecast back in as the next day’s lag, repeat. Simple, but errors compound.
- Direct, one model per horizon. The model for horizon h only uses lags of h days or more. Twenty-eight models for a 28-day horizon.
- Direct, one pooled model. Compute every feature at the origin, add the horizon as a feature, and train on every (origin, horizon) pair in the history.
My project uses the third. For a target ten days ahead, the “same weekday last week” feature has to be 14 days before the target, because 7 days before it is still in the future:
back = target - WEEK * np.ceil(horizon / WEEK).astype(int)
same_weekday = (y[back] + y[back - 7] + y[back - 14] + y[back - 21]) / 4
One guard makes this safe by construction: at forecast time the array y ends at the origin. A feature that reached past it would raise an IndexError instead of leaking quietly.
Calendar features are different. Bank holidays are published years ahead, so a flag for a future bank holiday is fair. Weather is not, unless you use the forecast that existed at the origin rather than the temperature that was later observed.
Common mistakes
- Features built on the full series. Any
shift,rollingorewmcomputed before the split can see the future for horizons beyond one. Compute features inside the loop, from the truncated history. - Imputation or scaling with full-series statistics. Filling gaps with the series mean, or standardising with the overall standard deviation, leaks a little of the future into every origin.
- Tuning on the test origins. If you choose features by their score on the same origins you report, the report is optimistic. In the project, 2024’s origins were for choices and interval calibration, and 2025’s were for scoring.
- Averaging away the horizon. Report error by horizon. In the project, exponential smoothing beat gradient boosting in the first week (MAE 25.1 against 25.8) and lost in each of the next three.
- Too few origins, or all in one season. A winter-only backtest says little about August bank holidays.
- Scoring days that didn’t happen. Outage days, or days filled by your own imputation, should not count as actuals.
- Using today’s data for yesterday’s forecast. Emergency department records can arrive late or be corrected. If the last few days before an origin would have been incomplete at the time, the backtest should see them incomplete too.
- Refitting differently from production. If the live model is retrained monthly, refit at each monthly origin. If it’s frozen for a year, freeze it in the backtest.
Once the backtest table exists, the next question is which numbers to compute from it. That’s the subject of forecast accuracy metrics. For forecasts that have to add up, such as specialties to a hospital total, hierarchical forecast reconciliation backtests the reconciled numbers the same way.
See it in a project
Tags
- backtesting
- time-series
- cross-validation
- leakage
- python