30-day readmission risk model
A readmission risk score for a discharge team that can only call so many patients, tested on a time split for calibration, subgroup performance and a planted data leak.
Synthetic data Every record here is generated. No real patient or organisational data is used.
- Test-period AUROC of gradient boosting (logistic regression 0.744)
- 0.753
- Readmissions reached by calling the top 10% of discharges
- 34%
- Observed against predicted readmission rate in the test period
- 11.6% vs 13.0%
- AUROC with a leaky column left in
- 0.845
A discharge team can phone a limited number of patients in the days after they go home. This project builds a 30-day emergency readmission risk score to decide who gets those calls, then checks the things that decide whether a score like this is safe to use: whether its probabilities can be taken at face value, how it behaves across patient groups, and whether any column is quietly handing it the answer.
All the data is synthetic, generated by the project code with a fixed seed for a fictional acute hospital. No real patient, hospital or staff data is involved.
In short: gradient boosting reached an AUROC of 0.753 on the most recent six months, 0.008 above logistic regression (95% bootstrap interval 0.002 to 0.015). Calling the top 10% of discharges would reach 34% of readmissions with either model, against 30% for a simple rule based on prior admissions. Both models over-predicted in the test period, 13.0% against an observed 11.6%, because readmissions had been falling. And given every column in the extract, boosting scored 0.845, thanks to a column that had been overwritten after the fact.
The problem
The team’s capacity is fixed, so the score’s first job is ranking: put the patients most likely to come back at the top of the list. Its second job is counting. A manager planning the service, or judging it a year later, needs to know how many readmissions to expect among the people called. That needs probabilities that match reality, not just a good order.
In the test period the hospital discharged 6,502 adults in six months. A call list of the top 10% is about 108 calls a month. So the questions that matter are concrete: how many readmissions do those 108 calls reach, what share of the people called would have been readmitted, and who never makes the list?
Data and what’s planted
The simulation produces 40,198 discharges for 25,189 patients between January 2023 and December 2025, with 12.8% readmitted as an emergency within 30 days (13.3% in 2023, 12.8% in 2024, 12.2% in 2025).
Each synthetic patient has a stream of admissions from 2021, so “emergency admissions in the prior 12 months” is a real count from their simulated history rather than a random number. When a readmission is drawn, it becomes the patient’s next admission, so the rows agree with each other: a discharge flagged as readmitted is followed by an emergency admission for the same patient within 30 days.
| Column | What it holds |
|---|---|
age, sex, deprivation_quintile |
Deprivation 1 is most deprived; missing for 2.2% of discharges |
admission_method, specialty |
Elective or emergency; eight specialties |
length_of_stay_days |
Skewed; same-day discharges are 0 |
emergency_admissions_12m |
Emergency admissions in the 365 days before this one |
comorbidity_count |
Coded comorbidities, capped at 10 |
frailty_flag |
Positive frailty screen; screening applies from 65 and is unrecorded for 9.4% of eligible stays |
discharge_weekday, discharge_destination |
Fewer weekend discharges; home, care package, care home, community hospital, self-discharge |
ward_code |
The trap (see below) |
The true readmission model is a logistic model with some things a plain logistic regression can’t see:
- non-linear effects of age and length of stay, including a bump for very short emergency stays;
- interactions: frail patients discharged home without support, frail patients discharged from Friday to Sunday, adults under 45 in the most deprived quintile, frequent attenders with several comorbidities, and prior admissions counting for more under 50;
- an unobserved patient-level vulnerability and per-stay noise, which cap what any model can reach;
- a downward trend in readmission risk and a small winter rise;
- one leak.
ward_codelooks like the discharge ward. It is actually the ward of the patient’s most recent admission at the extract date, joined from a patient-level snapshot. For anyone readmitted, it has been overwritten by the ward of the readmission or a later stay.
The generating probabilities, which no real extract would have, give an AUROC of 0.816 on the test period. That’s the ceiling with perfect knowledge of the planted model, including the unobserved parts.
Approach
A split that respects time
A score like this is used on next month’s patients, so it has to be tested on discharges later than anything it learned from.
# Each period ends a month before the next starts, so no 30-day outcome window crosses a boundary.
PERIODS = {
"train": ("2023-01-01", "2024-11-30"),
"validation": ("2025-01-01", "2025-05-31"),
"test": ("2025-07-01", "2025-12-31"),
}
That gives 25,573 training discharges (3,355 readmissions), 5,811 for validation (748) and 6,502 for test (755). The gap months matter. Training stops at the end of November, so its last outcomes are settled by the end of December. If it ran to the end of December, some of its labels would depend on readmissions in January 2025, inside the next period.
The five validation months chose the boosting settings and fitted the recalibration maps. The test half-year played no part in fitting, tuning or recalibration.
Two models
Logistic regression, with preprocessing chosen for this data rather than defaults:
def logistic_pipeline() -> Pipeline:
"""Splines for age, logs for skewed counts, one-hot categoricals, L2 penalty (C=1)."""
log_scale = make_pipeline(
FunctionTransformer(np.log1p, feature_names_out="one-to-one"), StandardScaler()
)
preprocess = ColumnTransformer(
[
("age", SplineTransformer(n_knots=5, degree=3), ["age"]),
("counts", log_scale, ["length_of_stay_days", "emergency_admissions_12m"]),
("comorbidities", StandardScaler(), ["comorbidity_count"]),
("categories", OneHotEncoder(handle_unknown="ignore"), CATEGORICAL),
]
)
return make_pipeline(preprocess, LogisticRegression(max_iter=2000))
And scikit-learn’s HistGradientBoostingClassifier, reading the categorical columns natively. Three settings were compared on validation log loss; the winner used shallow trees (at most 4 leaves), a learning rate of 0.05 and 500 iterations. Missing deprivation and unrecorded frailty get explicit levels (“Unknown”, “Not recorded”) for both models, so neither model imputes them silently and both treat them the same way.
For a baseline anyone could run without a model, I ranked patients by their prior emergency admissions, breaking ties at random.
Results
On the test half-year:
| Logistic regression | Gradient boosting | |
|---|---|---|
| AUROC (95% interval) | 0.744 (0.726 to 0.762) | 0.753 (0.735 to 0.772) |
| Average precision | 0.356 | 0.360 |
| Brier score | 0.0899 | 0.0897 |
| Scaled Brier (skill over the base rate) | 0.124 | 0.126 |
| Calibration intercept | −0.15 | −0.15 |
| Calibration slope | 0.94 | 0.95 |
| Mean predicted risk (observed 11.6%) | 13.0% | 13.0% |
Boosting’s AUROC gain is 0.008, with a paired bootstrap interval of 0.002 to 0.015. The interval excludes zero, the interactions I planted are the likely source, and the gain is small. The simple prior-admissions rule reached an AUROC of 0.655.
Data table
| Mean predicted risk | Logistic regression | Gradient boosting |
|---|---|---|
| 0.025 | 1.7% | |
| 0.026 | 2.6% | |
| 0.037 | 3.8% | |
| 0.04 | 3.7% | |
| 0.051 | 5.5% | |
| 0.052 | 4.9% | |
| 0.064 | 5.4% | |
| 0.066 | 6.8% | |
| 0.078 | 7.4% | |
| 0.08 | 8.0% | |
| 0.094 | 8.9% | |
| 0.097 | 9.1% | |
| 0.118 | 11.5% | |
| 0.12 | 11.4% | |
| 0.154 | 11.4% | |
| 0.155 | 11.4% | |
| 0.22 | 18.5% | |
| 0.226 | 20.6% | |
| 0.446 | 39.8% | |
| 0.448 | 39.8% |
Synthetic data; test period July to December 2025. Source: projects/readmission-risk-model.
Both models track the diagonal in the lower tenths and sit below it in the three highest-risk tenths: in the top tenth, predicted risk averaged about 45% and 39.8% were readmitted. Overall they predicted 13.0% against an observed 11.6%, a calibration intercept of −0.15. The slopes of 0.94 and 0.95 say the spread of the predictions is about right.
The over-prediction is drift, not overfitting. The training period’s readmission rate was 13.1%, validation’s 12.9% and the test period’s 11.6%. A Platt recalibration fitted on the validation months pulled the mean prediction down to 12.6% and the intercept to −0.11, a partial fix, because January to May was not yet as low as July to December. Calibration before AUC covers what these numbers mean and how to fix them.
What the discharge team would get
Data table
| Share of discharges called | Logistic regression | Gradient boosting | Simple rule: prior emergency admissions |
|---|---|---|---|
| 0 | 0.0% | 0.0% | 0.0% |
| 0.02 | 11.0% | 11.1% | 9.7% |
| 0.04 | 19.2% | 18.5% | 15.1% |
| 0.06 | 24.9% | 24.4% | 20.7% |
| 0.08 | 30.2% | 30.6% | 26.1% |
| 0.1 | 34.3% | 34.2% | 29.5% |
| 0.12 | 38.0% | 37.7% | 31.7% |
| 0.14 | 42.0% | 41.6% | 34.8% |
| 0.16 | 45.4% | 44.8% | 37.2% |
| 0.18 | 48.2% | 49.1% | 40.0% |
| 0.2 | 50.2% | 52.1% | 42.8% |
| 0.22 | 52.2% | 54.8% | 46.2% |
| 0.24 | 54.4% | 56.8% | 48.9% |
| 0.26 | 56.0% | 58.5% | 50.3% |
| 0.28 | 58.0% | 60.1% | 51.5% |
| 0.3 | 60.0% | 61.9% | 52.7% |
| 0.32 | 62.0% | 63.4% | 53.5% |
| 0.34 | 64.5% | 66.1% | 55.1% |
| 0.36 | 66.1% | 67.8% | 56.2% |
| 0.38 | 68.5% | 70.2% | 57.7% |
| 0.4 | 69.8% | 71.8% | 59.1% |
| 0.42 | 71.7% | 74.2% | 59.9% |
| 0.44 | 73.5% | 75.6% | 61.5% |
| 0.46 | 74.8% | 77.2% | 63.0% |
| 0.48 | 76.2% | 78.8% | 64.6% |
| 0.5 | 77.6% | 79.5% | 65.4% |
| 0.52 | 78.9% | 80.7% | 66.8% |
| 0.54 | 80.4% | 82.1% | 67.5% |
| 0.56 | 82.5% | 83.0% | 69.4% |
| 0.58 | 83.8% | 84.1% | 70.6% |
| 0.6 | 84.5% | 85.8% | 72.1% |
| 0.62 | 86.0% | 86.5% | 73.9% |
| 0.64 | 87.3% | 87.2% | 75.2% |
| 0.66 | 88.1% | 87.8% | 76.4% |
| 0.68 | 89.5% | 89.3% | 77.7% |
| 0.7 | 90.3% | 90.5% | 78.3% |
| 0.72 | 91.0% | 91.0% | 79.2% |
| 0.74 | 92.2% | 92.3% | 81.9% |
| 0.76 | 93.1% | 93.2% | 82.6% |
| 0.78 | 94.3% | 94.3% | 84.6% |
| 0.8 | 94.6% | 95.2% | 86.2% |
| 0.82 | 95.5% | 95.9% | 87.8% |
| 0.84 | 96.2% | 96.0% | 88.9% |
| 0.86 | 96.7% | 96.8% | 90.3% |
| 0.88 | 97.2% | 97.9% | 91.4% |
| 0.9 | 97.7% | 98.5% | 92.6% |
| 0.92 | 98.5% | 98.5% | 93.2% |
| 0.94 | 99.3% | 99.1% | 95.4% |
| 0.96 | 99.6% | 99.2% | 96.6% |
| 0.98 | 99.9% | 99.7% | 98.1% |
| 1 | 100.0% | 100.0% | 100.0% |
Synthetic data; test period July to December 2025. Source: projects/readmission-risk-model.
Readmissions reached out of 755 in the test half-year, with the share of all readmissions:
| Call list | Calls | Logistic regression | Boosting | Simple rule |
|---|---|---|---|---|
| Top 5% | 325 | 171 (22.6%) | 161 (21.3%) | 135 (17.9%) |
| Top 10% | 650 | 259 (34.3%) | 258 (34.2%) | 223 (29.5%) |
| Top 20% | 1,300 | 379 (50.2%) | 393 (52.1%) | 323 (42.8%) |
Positive predictive value, the share of people called who were readmitted:
| Call list | Logistic regression | Boosting | Simple rule |
|---|---|---|---|
| Top 5% | 52.6% | 49.5% | 41.5% |
| Top 10% | 39.8% | 39.7% | 34.3% |
| Top 20% | 29.2% | 30.2% | 24.8% |
At the team’s likely operating point, the top 10%, the two models reach almost the same number of readmissions: 259 and 258 of 755. Logistic regression does better on the shortest list and boosting on the longest. AUROC averages over every possible threshold; the team works at one. Both models beat the simple rule by 35 or 36 readmissions at 10%, and about four in ten of the people on a top-10% list were readmitted within 30 days, against 11.6% of all discharges.
Data table
| Column | Drop in AUROC |
|---|---|
| frailty_flag | 0.045 |
| admission_method | 0.037 |
| emergency_admissions_12m | 0.030 |
| comorbidity_count | 0.017 |
| age | 0.010 |
| deprivation_quintile | 0.009 |
| length_of_stay_days | 0.007 |
| specialty | 0.005 |
| discharge_destination | 0.003 |
| discharge_weekday | 0.000 |
| sex | -0.001 |
Synthetic data; test period July to December 2025. Source: projects/readmission-risk-model.
The boosting model leans on the frailty flag, admission method and prior emergency admissions. Discharge weekday barely registers, even though an interaction between frailty and Friday-to-Sunday discharge is planted: it is real but confined to a small group, and permutation importance measures average contribution across everyone. Explaining models to stakeholders shows how individual conditional expectation curves bring an interaction like this back into view.
Fairness checks
I checked discrimination, calibration and who ends up on the call list by deprivation quintile, age band and sex. The per-group numbers, with bootstrap intervals, are in the metrics download.
Data table
| Group | Logistic regression | Gradient boosting |
|---|---|---|
| Q1 (most deprived) | 0.727 | 0.746 |
| Q2 | 0.710 | 0.732 |
| Q3 | 0.764 | 0.766 |
| Q4 | 0.753 | 0.747 |
| Q5 (least deprived) | 0.761 | 0.749 |
| Age 18-44 | 0.696 | 0.727 |
| Age 45-64 | 0.646 | 0.677 |
| Age 65-79 | 0.713 | 0.714 |
| Age 80+ | 0.776 | 0.773 |
| Female | 0.747 | 0.751 |
| Male | 0.741 | 0.754 |
Synthetic data; test period July to December 2025. Source: projects/readmission-risk-model.
Discrimination is similar across deprivation quintiles and between sexes, and the per-group bootstrap intervals overlap widely (overlap isn’t a formal test of the difference, but nothing here suggests one): boosting’s AUROC runs from 0.732 (Q2, interval 0.693 to 0.775) to 0.766 (Q3, 0.715 to 0.814). It is weakest for patients aged 45 to 64 (0.677 for boosting, 0.646 for logistic regression), the group with the lowest readmission rate (7.2%) and without the frailty screen, which only applies from 65.
Data table
| Group | Logistic regression, predicted | Gradient boosting, predicted | Observed |
|---|---|---|---|
| Q1 (most deprived) | 16.1% | 15.8% | 14.4% |
| Q2 | 12.5% | 12.5% | 10.5% |
| Q3 | 11.5% | 11.6% | 9.5% |
| Q4 | 11.1% | 11.2% | 11.2% |
| Q5 (least deprived) | 9.9% | 10.1% | 9.1% |
| Age 18-44 | 11.0% | 10.9% | 9.9% |
| Age 45-64 | 7.3% | 6.9% | 7.2% |
| Age 65-79 | 10.7% | 11.0% | 9.2% |
| Age 80+ | 20.4% | 20.3% | 17.8% |
Synthetic data; test period July to December 2025. Source: projects/readmission-risk-model.
Calibration by group mostly repeats the overall over-prediction. Nothing stands out beyond what groups holding between 67 and 376 readmissions each can resolve.
The finding that matters is who gets called. A top-10% list by absolute risk goes mostly to the oldest patients:
- Patients aged 80 and over are 32.5% of discharges and 49.8% of readmissions, and they get 76.0% of the boosting model’s calls.
- Patients aged 45 to 64 have 15.8% of all readmissions and get 2 of the 650 calls. Their readmissions are effectively invisible to this service.
- The most deprived quintile gets 41.5% of calls for 38.5% of readmissions, in line with its higher rate (14.4% against 9.1% in the least deprived).
- Boosting reaches 24.1% of readmissions among 18 to 44 year olds, against 14.7% for logistic regression. That fits the planted interactions for younger patients (deprivation, and prior admissions counting for more under 50), which a logistic regression without interaction terms can’t represent.
None of this is the model misbehaving. It ranks by risk, and risk is concentrated in older patients. But “call the highest-risk patients” and “call the patients a call would help most” are different policies. If a call does more for a 55-year-old living alone than for a care-home resident with daily support, ranking by risk is the wrong objective, and fixing it needs evidence about the intervention, not a better model.
The leakage catch
Given every column in the extract, the boosting model scored an AUROC of 0.845 on the test period. For context, a systematic review of readmission models found C statistics of 0.68 to 0.83 for models usable at discharge (Kansagara et al., JAMA 2011). A model built from routine fields that beats the top of that range deserves suspicion before celebration. Here there’s a sharper test, because the data is synthetic: the true probabilities that generated the outcomes score 0.816, and no honest model can beat the truth. Scoring above it means the model is reading the answer.
A single-column screen didn’t find it. On its own, ward_code scored 0.575, eighth of twelve columns. Permutation importance did:
Data table
| Column | Drop in AUROC |
|---|---|
| ward_code | 0.112 |
| specialty | 0.078 |
| admission_method | 0.029 |
| frailty_flag | 0.027 |
| length_of_stay_days | 0.022 |
| emergency_admissions_12m | 0.018 |
| age | 0.007 |
| discharge_destination | 0.007 |
| comorbidity_count | 0.006 |
| deprivation_quintile | 0.005 |
| discharge_weekday | 0.000 |
| sex | 0.000 |
Synthetic data; test period July to December 2025. Source: projects/readmission-risk-model.
ward_code tops the list, and specialty jumps from 0.005 to 0.078. The two work as a pair: a surgical patient whose ward is a medical ward was admitted again after discharge. Two more checks confirmed it:
- Cross-tabulate against specialty. In the test period, 63.8% of readmitted discharges had a
ward_codebelonging to another specialty, against 5.2% of the rest. - Look within patients. For all 8,949 patients with more than one discharge,
ward_codeholds one value per patient, even though 84.5% of them were under more than one specialty. A stay-level field that never varies within a patient was joined at patient level.
The fix was to drop it. If the discharge ward is wanted, it has to come from the episode record as it stood at discharge. Data leakage in healthcare ML walks through this trap and a second one step by step.
Limits
- Synthetic data. I wrote the true model, so the results show how the method behaves, not how a real readmission model would perform. Real data would have messier coding and more of the missingness and drift patterns I only sketched.
- Not tested on real patients. One fictional hospital, one test half-year. This model is not fit for clinical use. A model built the same way on real data would still need external validation at another site and a later period before anyone relied on it.
- No deaths. A patient who dies after discharge can’t be readmitted. Counting them as “not readmitted” labels some of the sickest patients as good outcomes, which distorts both the model and its evaluation. Real work needs death as a competing outcome.
- Risk is not benefit. The model says who is likely to come back, not whose readmission a phone call would prevent.
What next
For a version built on real data, I’d do these next:
- Recalibrate on a rolling recent window and monitor observed against expected each month, since drift was the main calibration problem in this data.
- Report capture and positive predictive value at the team’s actual call capacity as the headline, with AUROC as a supporting number.
- Agree the ranking objective with the discharge team, including whether a separate list or quota for under-65s makes sense.
- Add a “known at discharge” column to the data dictionary and fail the pipeline when a model uses a field that isn’t.
Built with
- Python
- scikit-learn
- statsmodels
- pandas
- NumPy
Downloads
-
Synthetic discharges sample 2023-2025
19,998 discharges for whole patients chosen at random. ward_code is the planted leak; leave it out of any model.
-
Model metrics by subgroup
Test-period metrics in long format: one row per model, group and metric, including call-list capture and PPV.
- Project code