Behnam Analytics

Writing Forecasting

Forecast accuracy metrics and how to choose them

MAE, RMSE, MAPE, sMAPE, MASE, bias, interval coverage and pinball loss, what each one measures, where each misbehaves, and which to put in front of a planning audience.

Behnam Ebrahimi 8 min read

Every accuracy metric answers a different question. Pick the question first, then the metric, and the metric becomes easy to explain. Pick the metric first and you end up with a table of eight numbers and a planner asking which one matters.

The numbers below come from my emergency department demand forecasting project: four models, each forecasting daily attendances 28 days ahead from twelve monthly origins across 2025, with 335 scored days per model. The data is synthetic, at around 300 attendances a day.

Model MAE RMSE MAPE sMAPE MASE Bias 80% coverage Pinball
Seasonal naive 33.5 44.1 10.7% 10.4% 1.09 +2.0 74.3% 8.42
Exponential smoothing 24.9 32.6 7.9% 7.8% 0.81 +1.1 75.5% 6.36
Gradient boosting 24.2 31.2 7.6% 7.6% 0.79 −4.0 75.8% 5.87
Average of ETS and GBM 24.1 31.4 7.6% 7.5% 0.79 −1.4 76.4% 6.01

The average wins on MAE. Gradient boosting wins on RMSE and pinball loss. Exponential smoothing has the smallest bias. “Which model is best?” has no answer until you say what you are trying to minimise.

MAE

Mean absolute error is the average size of a miss, in the units of the forecast. “On a typical day the forecast is about 24 attendances out” is a sentence anyone in an operational meeting can use.

MAE treats a miss of 40 as exactly twice as bad as a miss of 20. The forecast that minimises it is the median of what might happen, so MAE is the natural choice when the cost of being wrong grows in proportion to the size of the miss, as it does for most staffing gaps.

Its weakness is that it has units. An MAE of 24 is good for a department seeing 300 a day and poor for one seeing 60, so you can’t compare across sites or specialties with it.

RMSE

Root mean squared error squares each miss before averaging, so large misses count for much more. Its optimal forecast is the mean rather than the median.

That makes RMSE the right choice when big misses are disproportionately expensive: an under-forecast that forces the hospital into escalation costs far more than two half-size misses. It also makes RMSE sensitive to a handful of days. In the project, every model missed Christmas Day by 85 attendances or more, and one day like that moves RMSE much more than MAE.

The ratio of the two tells you something. For normally distributed errors RMSE is about 1.25 times MAE (the square root of π/2). For all four models here it is about 1.3, so the errors have no unusually heavy tail. If that ratio jumps, look for a few very large misses before you look at anything else.

MAPE, and why it misbehaves

Mean absolute percentage error divides each miss by the actual. It’s popular because a percentage feels comparable across anything. It has three problems.

It explodes near zero. Two arrivals in a quiet overnight hour, forecast as four, is a 100% error. Three hundred in a day, forecast as 302, is 0.7%. Average those across a set of series and the quiet ones dominate. At zero it is undefined, and tools handle that differently. scikit-learn’s mean_absolute_percentage_error divides by a tiny epsilon instead, so mean_absolute_percentage_error([0, 100], [1, 100]) returns 2,251,799,813,685,248.

It’s asymmetric. With an actual of 100, a forecast of 150 is a 50% error. With an actual of 150, a forecast of 100 is 33%. The miss is the same 50 attendances; MAPE scores it differently depending on which side it falls. Under-forecasts can never cost more than 100%, while over-forecasts have no ceiling. If you choose between models by MAPE, you reward the ones that forecast low.

It weights days unevenly. A 20-attendance miss on a quiet Sunday counts for more than the same miss on a busy Monday, which is the opposite of what a planner usually cares about.

For this project MAPE behaves, because daily counts sit around 300 and never near zero: 7.6% for the best models. It would not behave for hourly arrivals, a small specialty, or a minor injuries unit’s overnight counts.

sMAPE

Symmetric MAPE divides by the average of the actual and the forecast rather than the actual alone, which caps it at 200% and removes the blow-up when the actual alone is zero. The name oversells it. With an actual of 100, a forecast of 150 scores 40% and a forecast of 50 scores 66.7%: the same 50-attendance miss, penalised more when it’s an under-forecast. Definitions also differ between sources, with and without the factor of two, so a sMAPE quoted without its formula is hard to compare.

MASE

Mean absolute scaled error divides your MAE by the in-sample MAE of a naive forecast. Hyndman and Koehler proposed it in 2006 to get a scale-free measure that doesn’t blow up near zero. For daily data with a weekly cycle, the natural naive is the seasonal one, “same day last week”:

MASE = MAE of the forecast / mean |y(t) − y(t − 7)| over the training data

Below 1 means the forecast beat what “same day last week” achieved in-sample. It works across series of different sizes, which makes it the metric I use to compare models or sites.

Two details trip people up. First, the denominator is computed on the training data, so in a rolling-origin backtest it changes with every origin; the project recomputes it on the history available at each one. Second, the denominator is a one-step naive: in-sample, “same day last week” always has last week’s actual to copy. Out of sample at 28 days it doesn’t, so the seasonal naive scores 1.09, not 1.00. Worse than 1 is normal for a naive forecast at long horizons, and it isn’t a bug.

MASE is hard to say out loud to a non-analyst. Use it to choose the model, then report MAE.

Bias

Bias is the average signed error. I define it as forecast minus actual, so a positive number means over-forecasting. State the sign convention every time, because both are in common use.

Bias matters more to a planner than its size suggests. Random error averages out over a rota; bias doesn’t. If the forecast is 20 low every day, you are short-staffed every day.

The annual figure can hide most of it. The average of ETS and GBM has a bias of −1.4 a day over 2025. Month by month it ranged from +25.3 in February to −28.8 in October:

Bias by forecast monthMean of forecast minus actual; above zero means over-forecasting
Data table
Forecast month, 2025Seasonal naiveExponential smoothingGradient boostingAverage of ETS and GBM
Jan4.57.3-1.62.8
Feb25.627.023.525.3
Mar-0.1-11.2-13.6-12.4
Apr5.42.5-0.11.2
May-13.5-3.4-17.5-10.4
Jun28.125.80.413.1
Jul-6.6-0.2-6.6-3.4
Aug-22.5-11.1-12.1-11.6
Sep10.42.33.73.0
Oct-16.6-30.3-27.4-28.8
Nov2.53.0-0.21.4
Dec7.11.43.02.2

Synthetic data. Source: projects/ed-demand-forecasting.

All four models moved together, which says the cause was in the data, not the models: a winter wave that faded after January, and a jump in level in October that nothing in the history predicted. Report bias by month, and by weekday if the plan is weekday-shaped. A cumulative sum of errors, checked each week, flags drift sooner than any annual table.

Interval coverage

If you publish an 80% interval, the actual should fall inside it on about 80% of days. Coverage is simply the share of days that did. In the project the four models covered 74.3% to 76.4%, so the intervals were a few points too narrow.

Coverage on its own can be gamed: an interval from zero to a thousand covers 100% of days and helps nobody. Always report it with the average width. The seasonal naive’s interval was 100.7 attendances wide on average; the average model’s was 72.3, with better coverage. Check coverage by horizon and by month too. An interval that is right on average but misses most of December isn’t right where it matters.

Pinball loss

Pinball loss, or quantile loss, scores a single quantile forecast. For the 90th percentile, a forecast that turns out too low costs 0.9 times the miss, and one that turns out too high costs 0.1 times it. The asymmetry is what pushes a forecast towards the right quantile:

def pinball(actual, quantile, tau):
    diff = actual - quantile
    return np.maximum(tau * diff, (tau - 1) * diff).mean()

scikit-learn has the same thing as mean_pinball_loss(y_true, y_pred, alpha=0.9).

Averaging the pinball loss of the 10th and 90th percentiles scores an 80% interval on coverage and width at once. That is why gradient boosting wins on pinball loss (5.87 against 6.01) even though the average had slightly better coverage: gradient boosting’s intervals were narrower, 70.5 against 72.3, and missed only a little more often. When you compare interval methods, compare them on pinball loss. When you explain them, use coverage and width.

The whole table in one function

With a backtest table of forecasts, actuals and interval bounds, every number above is a group-by:

def accuracy(df):
    actual, forecast = df["actual"], df["forecast"]
    error = forecast - actual
    return pd.Series(
        {
            "MAE": error.abs().mean(),
            "RMSE": np.sqrt((error**2).mean()),
            "MAPE": (error.abs() / actual).mean(),
            "sMAPE": (2 * error.abs() / (actual.abs() + forecast.abs())).mean(),
            "MASE": (error.abs() / df["scale"]).mean(),
            "bias": error.mean(),
            "coverage": actual.between(df["lower_80"], df["upper_80"]).mean(),
            "pinball": (pinball(actual, df["lower_80"], 0.1)
                        + pinball(actual, df["upper_80"], 0.9)) / 2,
        }
    )


scored.groupby("model").apply(accuracy)

Here scored holds the rows from the 2025 origins that have an actual, and scale is the in-sample seasonal naive MAE at each row’s origin. The project’s backtest forecasts file doesn’t include scale, so add it from the daily attendances file, using the history up to each origin. With that column, the function reproduces the table at the top.

Choosing for a planning audience

This is what I’d put in front of people who staff a department:

  1. MAE in their units, with the baseline next to it: “about 24 attendances a day out, 28% better than repeating last week”.
  2. Bias by month, in words: “we under-forecast October by about 29 a day”. This is the number that changes a rota.
  3. Interval coverage and width: “the 80% range held on 76% of days, and it’s about 72 attendances wide”.
  4. The days where it fails, named: Christmas, bank holidays, heatwaves.

And behind the scenes:

  • MASE to choose between models and to compare sites of different sizes.
  • Pinball loss to choose between interval methods.
  • RMSE when large misses cost disproportionately more.
  • MAPE only for series well away from zero, and never averaged across series of very different sizes.

Whatever you choose, choose it before you look at the results. A metric picked after the fact is usually the one your favourite model happened to win on. For how to produce the backtest table these metrics come from, see backtesting forecasts with rolling origins; for building and presenting intervals, see prediction intervals for planners.