Behnam Analytics

Writing Machine learning

Data leakage in healthcare ML, and how to catch it

Where leakage hides in health data, from post-discharge fields to time travel in lag features, how to detect it, and two worked examples from a readmission model.

Behnam Ebrahimi 8 min read

Leakage is information in the training data that won’t exist when the model is used. In health data it’s rarely a coding bug. It’s usually a field whose meaning depends on when you looked at it. The symptom is almost always a score that is better than it should be, and the question that finds it is almost always the same: when would I have known this?

The worked examples come from my 30-day readmission risk project, which runs on synthetic data for a fictional hospital. I planted one leak in the extract and built a second one on purpose as a feature. With the honest columns, gradient boosting scored an AUROC of 0.753 on the test period. With the planted leak it scored 0.845. With the engineered one, 0.909.

Where leakage hides in health data

Fields recorded after the event

A readmission model is used at discharge, so anything recorded afterwards is off limits: follow-up call outcomes, outpatient attendances after discharge, community referrals made later, anything with “latest” or “current” in its name. The dangerous ones don’t look post-discharge. Patient master tables often hold a single current value (ward, GP practice, risk score, deceased flag) that is overwritten as things change. Joined onto historical rows, that value describes the patient at extract time, not at discharge.

Coding done after the event

Diagnosis and comorbidity coding is usually completed after discharge, and records can be revised later. A model trained on the final, fully coded record expects a level of detail it won’t have when it scores a patient on the day they leave, and a revision made after a readmission can carry information about it. If the model runs at discharge, train it on the fields as they stood at discharge, or on what is reliably available by then.

Patient overlap across splits

A random row split scatters a patient’s discharges across training and test. In the project data, a random 80/20 split leaves 53.7% of test discharges with another discharge for the same patient in the training set. The model gets to learn from a patient’s later stays and is then scored on their earlier ones. Anything that nearly identifies a person (a rare combination of age, postcode and conditions) lets it recognise them.

Split by patient (GroupKFold, StratifiedGroupKFold or GroupShuffleSplit with patient ID as the group) or by time. A time split still has repeat patients (49.1% of the project’s test discharges belong to patients seen earlier), but that’s honest: the earlier stays really are in the past, which is how the model will meet returning patients in use.

Target-derived features

Any feature computed from the outcome, or from data that includes it: “readmissions this year”, a specialty’s readmission rate calculated over the whole dataset, a patient’s historical readmission rate that counts the current stay. They often come from BI tables built for reporting, where using the full year is correct.

Preprocessing fitted on the full data

Fitting a scaler on all the data before splitting is a small leak. Fitting anything that looks at the outcome is a large one: target encoders, univariate feature selection, outcome-aware imputation, and oversampling such as SMOTE before cross-validation, which creates synthetic training rows interpolated from patients who then land in the test fold. Put every step in a scikit-learn Pipeline so cross-validation refits it inside each fold. Resamplers such as SMOTE need imbalanced-learn’s Pipeline, because scikit-learn’s doesn’t accept them as steps. TargetEncoder helps here too: its fit_transform uses internal cross fitting, so a row’s encoding never includes its own outcome.

Time travel in lags and windows

Rolling counts that include the index date or run past it. Outcome windows that aren’t complete at the extract date, so recent discharges look like non-readmissions because there hasn’t been time. And split boundaries: a discharge on 20 December has an outcome decided in January, so if training ends in December and testing starts in January, training labels depend on test-period events. I leave a one-month gap between periods. TimeSeriesSplit has a gap argument, but it counts rows, not days, so for a 30-day embargo on irregular dates I define the periods by date.

How to detect it

  • Too-good results. Know the plausible range before you look at your own number. A systematic review of readmission models found C statistics of 0.68 to 0.83 for models usable at discharge (Kansagara et al., JAMA 2011). A big jump from adding one field is the clearest signal of all.
  • Single-feature checks. Score every column on its own. Any single field that rivals the whole model deserves a conversation with whoever owns it.
  • Importance on the full model. A column that dominates, or a second column that suddenly matters alongside it, points at the leak and its partner.
  • Time-based splits, and time-based diagnostics. Train on the past, test on the future, with a gap. Then check whether a feature’s relationship with the outcome changes with distance from the extract date. Honest features don’t care when the extract was taken.
  • A known-at audit. For every column, write down when it’s recorded and whether it’s ever updated. If nobody knows, that’s the finding.
  • Grain checks. A stay-level field should vary between a patient’s stays. If it never does, it was joined at patient level.

Worked example 1: the ward that moved

The extract has a column called ward_code that looks like the discharge ward. In the simulation it’s the ward of the patient’s most recent admission at the extract date, 1 February 2026, joined from a patient-level snapshot. For anyone readmitted, it has been overwritten by the ward of the readmission or a later stay.

The score. Gradient boosting with every column reached an AUROC of 0.845 on July to December 2025, against 0.753 without ward_code. Average precision went from 0.360 to 0.492. That’s above the top of the review’s range, from routine fields alone.

The single-feature screen missed it. I scored every column on its own over the training period, with cross-fitted target encoding for the categorical ones:

from sklearn.metrics import roc_auc_score
from sklearn.model_selection import StratifiedKFold
from sklearn.preprocessing import TargetEncoder

folds = StratifiedKFold(n_splits=5, shuffle=True, random_state=seed)
encoded = TargetEncoder(target_type="binary", cv=folds).fit_transform(train[["ward_code"]], y)
auc = roc_auc_score(y, encoded[:, 0])
How well each column predicts readmission on its ownCross-fitted AUROC on the training period; the axis starts at 0.5 (chance)
Data table
ColumnAUROC on its own
emergency_admissions_same_year0.900
emergency_admissions_12m0.654
comorbidity_count0.650
frailty_flag0.633
age0.613
admission_method0.608
length_of_stay_days0.603
specialty0.600
ward_code0.575
deprivation_quintile0.558
discharge_destination0.553
discharge_weekday0.522
sex0.510

Synthetic data; training period January 2023 to November 2024. Source: projects/readmission-risk-model.

ward_code scored 0.575 on its own, eighth of the twelve extract columns. A ward code alone says little, because a medical ward is normal for a medical patient. The leak lives in the combination.

Permutation importance found it. Shuffle one column at a time and measure the fall in test AUROC:

Permutation importance with ward_code left inMean fall in test-period AUROC when the column is shuffled (10 repeats)
Data table
ColumnDrop in AUROC
ward_code0.112
specialty0.078
admission_method0.029
frailty_flag0.027
length_of_stay_days0.022
emergency_admissions_12m0.018
age0.007
discharge_destination0.007
comorbidity_count0.006
deprivation_quintile0.005
discharge_weekday0.000
sex0.000

Synthetic data; test period July to December 2025. Source: projects/readmission-risk-model.

ward_code tops the list at 0.112, and specialty jumps to 0.078 from 0.005 in the honest model. They are working as a pair: a surgical patient whose ward is a medical ward has been admitted again since.

A cross-tab confirmed it. Using the ward list for each specialty, I flagged discharges whose ward_code belongs to a different specialty:

def outside_home_ward(frame: pd.DataFrame) -> pd.Series:
    """True where ward_code is not a ward the discharging specialty uses."""
    table = home_wards()
    pairs = zip(frame["specialty"], frame["ward_code"], strict=True)
    return pd.Series([w not in table[s] for s, w in pairs], index=frame.index)

In the test period, 63.8% of readmitted discharges had a ward from another specialty, against 5.2% of the rest. In the training period the figures were 69.9% and 31.8%. That difference is the time-travel signature:

Discharges whose ward_code belongs to a different specialtyThe gap between readmitted and other patients widens towards the extract date
Data table
Discharge quarterReadmitted within 30 daysNot readmitted
2023-01-0169%40%
2023-04-0175%38%
2023-07-0171%36%
2023-10-0171%33%
2024-01-0167%29%
2024-04-0171%27%
2024-07-0168%24%
2024-10-0166%21%
2025-01-0168%15%
2025-04-0165%13%
2025-07-0165%8%
2025-10-0163%3%

Synthetic data, all discharges 2023 to 2025; extract taken 1 February 2026. Source: projects/readmission-risk-model.

For patients who weren’t readmitted, the mismatch rate falls steadily towards the extract date. An older discharge has had longer for some later admission to overwrite the field; a recent one hasn’t. An honest discharge-ward field would look the same in every quarter. This one also means the leak is strongest in exactly the period used for testing.

The grain check closed it. Among the 8,949 patients with more than one discharge, ward_code has one value per patient every time, although 84.5% of them were under more than one specialty.

per_patient = frame.groupby("patient_id").agg(
    stays=("discharge_id", "size"),
    wards=("ward_code", "nunique"),
    specialties=("specialty", "nunique"),
)
per_patient = per_patient[per_patient["stays"] > 1]

The fix is to drop the column. If the discharge ward is wanted, it has to come from the episode record as it stood at discharge.

Worked example 2: the count that looked ahead

The honest model uses emergency admissions in the 365 days before the current admission. A reporting table might instead hold emergency admissions per patient per calendar year, which is a perfectly good BI measure. As a feature, it counts admissions later in the same year, including the readmission you are trying to predict:

def calendar_year_admissions(frame: pd.DataFrame) -> pd.Series:
    """The leaky way to count: every emergency admission in the discharge's calendar year."""
    admitted = frame["discharge_date"] - pd.to_timedelta(frame["length_of_stay_days"], unit="D")
    emergency = frame[frame["admission_method"] == "Emergency"]
    counts = emergency.groupby(["patient_id", admitted[emergency.index].dt.year]).size().rename("n")
    keys = pd.MultiIndex.from_arrays([frame["patient_id"], frame["discharge_date"].dt.year])
    return pd.Series(counts.reindex(keys).fillna(0).to_numpy(dtype=int), index=frame.index)

This one the single-feature screen catches at once: 0.900 on its own over the training period, against 0.654 for the honest prior-12-month count. On its own it outscores the whole honest model on the same period (0.785, even scored on the data it was trained on). Swapped into the model, it lifted test AUROC to 0.909. No real readmission model built at discharge should get near that.

The two examples are why I run more than one check. A screen catches a field that carries the answer by itself; importance and cross-tabs catch one that needs a partner; the time diagnostic and the grain check explain what happened, which is what you need before you can go back to the data’s owner and fix the source.

Checklist

  • Before modelling, write the prediction moment down (“at discharge”) and mark every column as known or not known by then.
  • Split by time with a gap at least as long as the outcome window, or by patient if time isn’t the point.
  • Keep every fitted step, including encoders and resampling, inside a pipeline.
  • Know the plausible range for the problem, and treat a result above it as a bug until proven otherwise.
  • Run a single-feature screen, then permutation importance on the full model, and look for pairs.
  • Check each suspicious field’s relationship with the outcome by extract distance, and its grain.
  • When you find a leak, fix it at the source, and add the check to the pipeline so it can’t come back.