Hierarchical forecast reconciliation
Forecasting specialty demand that has to add up to the hospital total. Why incoherent forecasts start planning arguments, what bottom-up and top-down get wrong, how MinT reconciliation works, and a backtest in Python.
Forecast the hospital’s emergency admissions, forecast each division, forecast each specialty, and the three answers won’t add up. Each model sees a different series, so each makes different mistakes. Reconciliation adjusts all the forecasts together so they add up, and it can make them more accurate while it does. In the backtest below, reconciliation cut the specialty-level error by about 3%. Top-down, the method many teams reach for first, raised it by a third.
The data is synthetic, generated by hierarchy_reconciliation.py with a fixed seed for a fictional acute hospital.
Why incoherent forecasts start arguments
Take the December 2024 origin of the backtest below. A separate model for each of the ten series in the hierarchy, summed over 2025, gives:
| Forecast for 2025 | Emergency admissions |
|---|---|
| Hospital total model | 48,274 |
| Sum of the two division models | 48,751 |
| Sum of the seven specialty models | 47,828 |
Three numbers for the same thing, 923 admissions apart, about 2.5 a day. The finance plan uses the total. Each division head brings their own forecast, and the two add up to more than the hospital’s. The specialty rotas are built from the specialty forecasts, which add up to less. The gap becomes a negotiation: whose number is right, and who absorbs the difference? Nobody’s number is right, and the argument is about modelling noise.
Data table
| Month | Actual | Hospital total model | Sum of specialty models | MinT (shrink), coherent |
|---|---|---|---|---|
| 2022-01-01 | 4,125 | |||
| 2022-02-01 | 3,594 | |||
| 2022-03-01 | 3,994 | |||
| 2022-04-01 | 3,634 | |||
| 2022-05-01 | 3,713 | |||
| 2022-06-01 | 3,542 | |||
| 2022-07-01 | 3,750 | |||
| 2022-08-01 | 3,759 | |||
| 2022-09-01 | 3,603 | |||
| 2022-10-01 | 3,804 | |||
| 2022-11-01 | 3,747 | |||
| 2022-12-01 | 4,021 | |||
| 2023-01-01 | 4,229 | |||
| 2023-02-01 | 3,751 | |||
| 2023-03-01 | 4,136 | |||
| 2023-04-01 | 3,821 | |||
| 2023-05-01 | 3,869 | |||
| 2023-06-01 | 3,670 | |||
| 2023-07-01 | 3,763 | |||
| 2023-08-01 | 3,638 | |||
| 2023-09-01 | 3,698 | |||
| 2023-10-01 | 3,874 | |||
| 2023-11-01 | 3,735 | |||
| 2023-12-01 | 4,202 | |||
| 2024-01-01 | 4,236 | |||
| 2024-02-01 | 4,012 | |||
| 2024-03-01 | 4,187 | |||
| 2024-04-01 | 3,884 | |||
| 2024-05-01 | 3,892 | |||
| 2024-06-01 | 3,665 | |||
| 2024-07-01 | 3,890 | |||
| 2024-08-01 | 3,847 | |||
| 2024-09-01 | 3,901 | |||
| 2024-10-01 | 3,880 | |||
| 2024-11-01 | 4,091 | |||
| 2024-12-01 | 4,131 | |||
| 2025-01-01 | 4,284 | 4,306 | 4,286 | 4,303 |
| 2025-02-01 | 3,992 | 3,881 | 3,861 | 3,878 |
| 2025-03-01 | 4,320 | 4,216 | 4,191 | 4,209 |
| 2025-04-01 | 3,923 | 3,947 | 3,915 | 3,934 |
| 2025-05-01 | 4,123 | 3,980 | 3,941 | 3,961 |
| 2025-06-01 | 3,809 | 3,822 | 3,779 | 3,798 |
| 2025-07-01 | 3,977 | 3,984 | 3,937 | 3,956 |
| 2025-08-01 | 3,803 | 3,965 | 3,919 | 3,939 |
| 2025-09-01 | 3,815 | 3,875 | 3,826 | 3,848 |
| 2025-10-01 | 4,185 | 4,038 | 3,994 | 4,015 |
| 2025-11-01 | 4,029 | 3,992 | 3,948 | 3,971 |
| 2025-12-01 | 4,577 | 4,268 | 4,231 | 4,253 |
Synthetic data. Source: projects/article-examples/round-two/hierarchy_reconciliation.py.
The chart also shows the MinT reconciled forecast (explained below), which puts 48,065 admissions in 2025 at every level. All of these were below the 48,837 that actually arrived; reconciliation makes forecasts agree, not clairvoyant.
Coherent forecasts end that argument before it starts. Every level gets a number, and the numbers add up. The only question left is how to make them add up without throwing away information.
The example
The hospital has two divisions and seven specialties:
- Medicine: general medicine, respiratory medicine, cardiology and geriatric medicine.
- Surgery: general surgery, trauma and orthopaedics, and urology.
Monthly emergency admissions run from January 2018 to December 2025, rising from 43,245 a year to 48,837. Planted in the data:
- Different trends. Cardiology and geriatric medicine grow, and general surgery shrinks as patients move to same-day emergency care. Between 2018 and 2025 cardiology’s share of admissions went from 11.1% to 15.8%, and general surgery’s from 15.1% to 9.8%.
- Different seasons. Respiratory medicine has a sharp winter peak. General medicine, geriatric medicine and cardiology have milder ones, and trauma and orthopaedics has a small summer bump.
- A shared winter. Winter severity changes from year to year and hits the medical specialties together, so their forecast errors are correlated.
- Month lengths, Poisson noise and a little extra wobble.
The backtest follows the method in backtesting forecasts with rolling origins: 25 monthly origins from December 2022 to December 2024, each forecasting the next 12 months with only the data available at the time. The base model for every series is a damped-trend Holt-Winters model on the log scale, fitted separately, much as a team would if each specialty owned its own forecast. I scored every method with the mean absolute scaled error (MASE), scaled by the in-sample error of a same-month-last-year forecast at each origin, and averaged over the series in each level. Below 1 means the forecast beat what the seasonal naive achieved in-sample. The emergency department demand forecasting project uses the same metric, and forecast accuracy metrics and how to choose them explains it.
Bottom-up and top-down
Bottom-up forecasts the specialties and adds them up. It keeps every specialty’s own trend and season, but the total inherits every specialty’s noise, and specialty series are noisier than the total.
Top-down forecasts the total and splits it by historical shares. Here, each specialty’s share of the last 12 months. The total is the smoothest series, so it’s usually the easiest to forecast. But fixed shares assume the mix doesn’t change, and they give every specialty the total’s seasonal shape. The fpp3 textbook notes a deeper problem: no top-down method keeps the forecasts unbiased, even when the base forecasts are.
Both showed their weaknesses:
| Level | Base | Bottom-up | Top-down | OLS | WLS (structural) | MinT (shrink) |
|---|---|---|---|---|---|---|
| Total | 0.895 | 0.966 | 0.895 | 0.890 | 0.899 | 0.904 |
| Division | 0.783 | 0.834 | 1.114 | 0.772 | 0.777 | 0.783 |
| Specialty | 0.777 | 0.777 | 1.038 | 0.749 | 0.752 | 0.756 |
Bottom-up’s total was 8% worse than the total model’s own forecast. Top-down’s specialties were 34% worse than their own models, with a MASE above 1. The damage was concentrated where I planted it. Respiratory medicine’s MASE doubled, from 0.927 to 1.951, because a sharp winter peak got the total’s gentle one. Cardiology and geriatric medicine got 40% worse, because shares from last year lag a growing specialty, and general surgery got 20% worse, because they overstate a shrinking one. On average, top-down under-forecast cardiology by 32.6 admissions a month, against 7.6 for its own model.
Reconciliation as a projection
Reconciliation uses all the base forecasts at once. Write the hierarchy as a summing matrix S, which maps the seven specialty values to all ten series:
GM Resp Card Geri GS T&O Uro
Hospital total 1 1 1 1 1 1 1
Medicine 1 1 1 1 0 0 0
Surgery 0 0 0 0 1 1 1
General medicine 1 0 0 0 0 0 0
... one row per specialty: the 7 × 7 identity
Any coherent set of forecasts is S times some vector of specialty forecasts. A linear reconciliation method picks a matrix G that turns the ten base forecasts ŷ into seven specialty forecasts, and then adds them up: ỹ = S G ŷ. Bottom-up is a G that ignores everything except the specialty rows. Top-down is a G that ignores everything except the total.
MinT (minimum trace), from Wickramasuriya, Athanasopoulos and Hyndman (2019), chooses G by asking two things. First, if the base forecasts are unbiased, the reconciled ones should be too; the condition is SGS = S. Second, among all such G, pick the one that minimises the total variance of the reconciled forecast errors across every series. The answer is:
G = (S′ W⁻¹ S)⁻¹ S′ W⁻¹
where W is the covariance matrix of the base forecast errors. It’s a generalised least-squares fit of the coherent structure to the base forecasts: series with noisy forecasts get less weight, and errors that move together are accounted for. In numpy it’s three lines:
def reconcile(S: np.ndarray, base: np.ndarray, W: np.ndarray) -> np.ndarray:
"""y_tilde = S (S' W^-1 S)^-1 S' W^-1 y_hat, for base forecasts in columns."""
Winv = np.linalg.inv(W)
G = np.linalg.solve(S.T @ Winv @ S, S.T @ Winv)
return S @ G @ base
The hard part is W, which you never know. The standard choices, as fpp3 describes them, differ only in what they assume about it:
- OLS: W is the identity. Every series’ errors are equally large and uncorrelated.
- WLS (structural): W is diagonal, proportional to the number of specialties under each node, so the total counts for seven and each specialty for one.
- MinT (shrink): W is the covariance of each model’s one-step in-sample residuals, shrunk towards its diagonal because ten series and a few years of data can’t estimate 45 correlations well. The shrinkage intensity is estimated from the data; here it was 0.18 at the median origin.
In the script, the three are an identity matrix, np.diag(S.sum(axis=1)) and a short shrinkage estimator. If you use R, fabletools::min_trace() offers these as method = "ols", "wls_struct" and "mint_shrink". In Python, Nixtla’s hierarchicalforecast package has a MinTrace reconciler.
Results
Data table
| Level | Base | Bottom-up | Top-down | OLS | WLS (structural) | MinT (shrink) |
|---|---|---|---|---|---|---|
| Total | 0.90 | 0.97 | 0.90 | 0.89 | 0.90 | 0.90 |
| Division | 0.78 | 0.83 | 1.11 | 0.77 | 0.78 | 0.78 |
| Specialty | 0.78 | 0.78 | 1.04 | 0.75 | 0.75 | 0.76 |
Synthetic data. Source: projects/article-examples/round-two/hierarchy_reconciliation.py.
All three reconciliation methods improved the specialty forecasts: OLS by 3.6%, WLS by 3.2% and MinT by 2.7%. They left the total roughly where it was, between 0.6% better and 1.0% worse. And they did it while making every level add up, which bottom-up and top-down also do, at a much higher price.
Data table
| Specialty | Change in MASE |
|---|---|
| General medicine | -12% |
| Respiratory medicine | 110% |
| Cardiology | 40% |
| Geriatric medicine | 40% |
| General surgery | 20% |
| Trauma and orthopaedics | 17% |
| Urology | 7% |
| Specialty | Change in MASE |
|---|---|
| General medicine | -2% |
| Respiratory medicine | -3% |
| Cardiology | -5% |
| Geriatric medicine | -11% |
| General surgery | -5% |
| Trauma and orthopaedics | -0% |
| Urology | 2% |
| Specialty | Change in MASE |
|---|---|
| General medicine | -1% |
| Respiratory medicine | -3% |
| Cardiology | -4% |
| Geriatric medicine | -8% |
| General surgery | -4% |
| Trauma and orthopaedics | -2% |
| Urology | -1% |
| Specialty | Change in MASE |
|---|---|
| General medicine | -2% |
| Respiratory medicine | -2% |
| Cardiology | -2% |
| Geriatric medicine | -6% |
| General surgery | -5% |
| Trauma and orthopaedics | -2% |
| Urology | -2% |
Synthetic data. Source: projects/article-examples/round-two/hierarchy_reconciliation.py.
Switch between the methods to see where each one helped. The reconciliation methods improved six or seven of the seven specialties each, with the largest gains in geriatric medicine, where OLS cut MASE by 11%. Top-down made six of seven worse. The exception was general medicine, 12% better. It’s the largest specialty, so its seasonal shape is close to the total’s, which is probably why the smoother total forecast helped it more than the fixed share hurt.
MinT didn’t win here. In practice it estimates W from one-step in-sample residuals and assumes errors further ahead share their structure, up to a scale factor. With ten series and a few years of monthly data, that estimate is noisy, and the simpler weights did as well. I’d read the three as equivalent on this data and pick whichever is easiest to explain.
A caveat on reading the table: MASE was averaged over series within each level, so the specialty row treats urology (about 2,100 admissions a year) the same as general medicine (about 14,600). If the plan cares about admissions rather than specialties, weight by volume, or report absolute errors too.
Which to use
- Don’t use fixed-share top-down when the mix is changing or when the lower levels have their own seasonal shapes. Both are normal in hospital activity.
- Bottom-up is fine when the bottom level is large and well-behaved. Here it cost 8% at the total.
- Start with OLS or WLS (structural). They need no residuals, can’t be destabilised by a bad covariance estimate, and did as well as MinT here.
- Try MinT (shrink) when you have many series with correlated errors and enough history to estimate them, and backtest it against the simpler weights rather than assuming it wins.
- Reconcile the forecasts you’ll publish, then backtest the reconciled ones. The accuracy that matters is the accuracy of the numbers people plan with.
Reconciliation doesn’t fix a bad base model. Every method here under-forecast cardiology and geriatric medicine on average, most likely because the damped trend in the base model assumed their growth would slow, and it didn’t. Reconciliation shares information across the hierarchy, and if every series has the same blind spot, the reconciled forecasts inherit it. For the intervals that go with the point forecasts, see prediction intervals for planners.
Reproduce
The data, the backtest, the tables and the three charts come from one script. It fits 250 Holt-Winters models, so it takes a couple of minutes:
uv run python projects/article-examples/round-two/hierarchy_reconciliation.py
See it in a project
Tags
- forecasting
- hierarchical-forecasting
- reconciliation
- mint
- backtesting
- python