Behnam Analytics

Writing Machine learning

Calibration before AUC

Why a risk score used for decisions needs probabilities that match reality, and how to check and fix calibration in scikit-learn, with numbers from a readmission model.

Behnam Ebrahimi 8 min read

A readmission model can rank patients well and still put the average patient’s risk at 41% when the real rate is nearer 12%. AUROC will not notice. If anyone reads the number itself, to set a threshold, plan a service or talk to a patient, calibration matters as much as ranking, and it breaks more easily.

The numbers below come from my 30-day readmission risk project, which runs on synthetic data for a fictional hospital: models trained on discharges from January 2023 to November 2024, tuned and recalibrated on January to May 2025, and tested on July to December 2025 (6,502 discharges, 11.6% readmitted).

What AUROC does and doesn’t tell you

AUROC is the probability that a randomly chosen patient who was readmitted gets a higher score than a randomly chosen patient who wasn’t. Only the order of the scores matters. Double every score, square it, add a constant to the log-odds: the order doesn’t change, so neither does AUROC.

That makes AUROC a good summary of ranking and a blind one for everything else. It can’t tell you:

  • whether a predicted 20% means about one in five;
  • how many events to expect among the patients you act on;
  • whether a threshold like “everyone above 30%” picks out 50 patients or 2,000;
  • whether the score still means the same thing six months later.

It doesn’t even tell you much about a ranking decision at a fixed capacity. In the project, gradient boosting beat logistic regression on AUROC by 0.008. At the discharge team’s operating point, calling the top 10% of patients, logistic regression reached 259 readmissions and boosting 258. AUROC averages over every possible threshold; a service works at one.

So when does calibration matter? Whenever the probability leaves the model as a number: thresholds on absolute risk, expected counts for planning, shared decisions with patients, comparing risk across sites or combining it with costs. For a clinical or operational risk score, that is most of the time.

The class-weights trap

One of the most common ways to break calibration is deliberate. The outcome is rare, so someone reaches for class weights or resampling to “fix the imbalance”:

from sklearn.ensemble import HistGradientBoostingClassifier

model = HistGradientBoostingClassifier(class_weight="balanced", random_state=0)

"balanced" weights each class inversely to its frequency, so the model trains as if readmissions were as common as non-readmissions. It learns roughly the same ranking and a very different scale. On the test period:

Unweighted Balanced class weights
AUROC 0.753 0.753
Mean predicted risk (observed 11.6%) 13.0% 41.4%
Brier score 0.0897 0.1915
Scaled Brier 0.126 −0.866
Calibration intercept −0.15 −2.00
Calibration slope 0.95 0.97

The AUROCs agree to three decimal places. The weighted model’s average prediction is more than three times the real rate, and its scaled Brier score is negative: it is worse than predicting 11.6% for everyone. How much harm weights do depends on how rare the outcome is: in the referral triage text classifier, where the classes are far less lopsided, balanced weights barely moved the Brier score (0.0303 without, 0.0290 with).

Same ranking, different probabilitiesGradient boosting trained with and without balanced class weights, by tenth of risk
Data table
Mean predicted riskGradient boostingGradient boosting, balanced class weights
0.0251.7%
0.0373.8%
0.0515.5%
0.0645.4%
0.0787.4%
0.0948.9%
0.11811.5%
0.131.7%
0.15511.4%
0.193.8%
0.22620.6%
0.2556.2%
0.3075.8%
0.3576.0%
0.4069.4%
0.44839.8%
0.4711.1%
0.54811.4%
0.65521.2%
0.8239.5%

Synthetic data; test period July to December 2025. Source: projects/readmission-risk-model.

Its lowest-risk tenth of patients was told 13% on average, and 1.7% were readmitted. Its highest-risk tenth was told 82%, and 39.5% were.

This isn’t a quirk of one library. Van den Goorbergh and colleagues (JAMIA, 2022) tested random undersampling, random oversampling and SMOTE with standard and ridge logistic regression and found the same pattern: strong overestimation of risk, no better AUROC. If you need more positives flagged, lower the threshold on a calibrated model instead.

Checking calibration

The calibration plot

Sort patients by predicted risk, split them into groups, and plot each group’s mean prediction against the share who actually had the outcome. Perfect calibration sits on the diagonal.

from sklearn.calibration import calibration_curve

observed, predicted = calibration_curve(y_test, p_test, n_bins=10, strategy="quantile")

Use strategy="quantile" for rare outcomes. The default, "uniform", uses equal-width bins, which leaves the high-risk bins almost empty when most predictions sit below 20%. Quantile bins put the same number of patients in each point, so no point rests on a handful of patients.

Always draw it on data the model hasn’t seen, and ideally on a later period than it was trained on. A model is usually well calibrated on its own training data; that’s the least interesting place to look.

Calibration intercept and slope

The plot is the thing to look at. Two numbers summarise it for tables and monitoring:

  • Calibration intercept (calibration-in-the-large). Fit a logistic regression of the outcome with the model’s log-odds as a fixed offset and only an intercept free. Zero is ideal. Negative means the model over-predicts on average; positive means it under-predicts.
  • Calibration slope. Regress the outcome on the model’s log-odds. One is ideal. Below one, predictions are too extreme (high risks too high, low risks too low), the usual sign of overfitting. Above one, they’re too timid.

scikit-learn doesn’t compute these, so I use statsmodels:

import numpy as np
import statsmodels.api as sm


def calibration_intercept_slope(y: np.ndarray, p: np.ndarray) -> tuple[float, float]:
    """Calibration-in-the-large (logit p as an offset) and calibration slope (on logit p)."""
    p = np.clip(p, 1e-6, 1 - 1e-6)
    lp = np.log(p / (1 - p))
    binomial = sm.families.Binomial()
    intercept = sm.GLM(y, np.ones((len(y), 1)), family=binomial, offset=lp).fit().params[0]
    slope = sm.GLM(y, sm.add_constant(lp), family=binomial).fit().params[1]
    return float(intercept), float(slope)

In the project, both models had test-period slopes close to one (0.94 for logistic regression, 0.95 for boosting) and intercepts of −0.15: the spread of risk was right, the level was too high. Van Calster and colleagues (BMC Medicine, 2019) set out this hierarchy of calibration checks in more depth, and it’s the paper I’d hand to anyone new to the topic.

Brier score

The Brier score is the mean squared difference between the predicted probability and what happened (1 or 0). Lower is better, and it rewards both ranking and calibration. On its own it’s hard to read, because it depends on the base rate, so compare it with the Brier score of predicting the base rate for everyone:

from sklearn.metrics import brier_score_loss, d2_brier_score

brier = brier_score_loss(y_test, p_test)
scaled = d2_brier_score(y_test, p_test)  # 1 - Brier / Brier of always predicting the base rate

scikit-learn calls the scaled version d2_brier_score; it’s also known as the Brier skill score. Boosting scored 0.0897 (scaled 0.126) and logistic regression 0.0899 (0.124). Scores like these say the probabilities carry real information beyond the base rate, not that they’re precise for any one patient.

Recalibration

If the ranking is good and the scale is wrong, keep the model and learn a map from its scores to probabilities, using data it wasn’t trained on. Two standard maps:

  • Platt scaling fits a logistic regression to the model’s score: two parameters, which correct the intercept and slope. It can’t fix a curve that bends in the middle, and it needs little data.
  • Isotonic regression fits a non-decreasing step function. It can fix any monotone distortion but needs more data. scikit-learn’s documentation advises against it with well under 1,000 calibration samples, and its flat steps tie patients who had different scores.

For a model that is already fitted, wrap it in FrozenEstimator so the calibrator doesn’t refit it:

from sklearn.calibration import CalibratedClassifierCV
from sklearn.frozen import FrozenEstimator

platt = CalibratedClassifierCV(FrozenEstimator(model), method="sigmoid")
platt.fit(X_calibration, y_calibration)  # held-out data: not the training set, not the test set
p_test = platt.predict_proba(X_test)[:, 1]

method="sigmoid" is Platt scaling. It works on the model’s decision_function when there is one, which for gradient boosting is the raw log-odds, and on predict_proba otherwise. Swap in method="isotonic" for the step function.

The calibration set matters as much as the method. I used the validation months, January to May 2025 (5,811 discharges, 748 readmissions): later than training, earlier than test. Recalibrating on the training data teaches the map nothing, because the model already fits that data too well; recalibrating on the test data leaves you nothing honest to report.

Applied to the class-weighted model:

Balanced weights After Platt After isotonic
AUROC 0.753 0.753 0.753
Mean predicted risk (observed 11.6%) 41.4% 12.6% 12.6%
Brier score 0.1915 0.0897 0.0899
Calibration intercept −2.00 −0.11 −0.11
Calibration slope 0.97 0.96 0.91
Average precision 0.355 0.355 0.333
Recalibrating the class-weighted model on held-out dataPlatt and isotonic maps fitted on January to May 2025, checked on the test period
Data table
Mean predicted riskGradient boosting, balanced class weightsBalanced, Platt scalingBalanced, isotonic regression
0.0141.7%
0.0221.7%
0.0343.8%
0.0455.3%
0.0496.2%
0.0625.8%
0.0646.0%
0.0776.0%
0.0785.8%
0.0939.4%
0.10510.6%
0.11411.3%
0.11711.1%
0.131.7%
0.15511.4%
0.16912.4%
0.193.8%
0.22421.2%
0.2521.6%
0.2556.2%
0.3075.8%
0.3576.0%
0.4069.4%
0.42539.6%
0.42839.5%
0.4711.1%
0.54811.4%
0.65521.2%
0.8239.5%

Synthetic data; test period July to December 2025. Source: projects/readmission-risk-model.

Platt scaling recovered almost everything and left the ranking untouched. Isotonic did about as well on Brier, but its steps cost some average precision (0.355 to 0.333) by tying patients together at the top of the list. With a few thousand calibration rows and a distortion that is basically a shift, the two-parameter map was the better choice.

Calibration drifts even when ranking holds

Both recalibrated models still show an intercept of −0.11. That’s not the method failing. Readmissions were falling: the observed rate was 13.1% in training, 12.9% in the validation months and 11.6% in the test months. The unweighted boosting model’s intercept moved from 0.00 on its training data to −0.04 on validation and −0.15 on test, while its AUROC went from 0.774 on validation to 0.753 on test.

A map fitted on January to May can only correct for January to May. Platt-recalibrating the unweighted model on those months moved its test-period mean prediction from 13.0% to 12.6%, part of the way to 11.6%. That’s the case for monitoring observed against expected every month once a score is live, and refitting the intercept on a recent window when the gap persists. In this data the ranking held up reasonably well; the level moved within six months.

What I report for any risk model

  • A calibration plot by tenth of predicted risk, on a period after the training data.
  • Calibration intercept and slope next to the plot, not instead of it.
  • Brier and scaled Brier alongside AUROC.
  • Capture and positive predictive value at the operating point the service will actually use.
  • Once live, observed against expected by month, with a rule for when to recalibrate.
  • No class weights or resampling in a model whose probabilities will be read, unless it’s recalibrated afterwards. Better still, leave them out.

The numbers belong in the model card that goes with the score; explaining models to stakeholders has a skeleton for one.