Logistic regression or gradient boosting for tabular health data
A head-to-head on synthetic readmission data with scikit-learn, comparing AUROC, Brier score, calibration and stability, and a decision guide for choosing between them.
For most risk models on tabular health data, the choice comes down to logistic regression or gradient boosted trees. Boosting usually wins the leaderboard. Logistic regression usually wins the meeting where someone has to explain, validate and maintain the model. The useful question isn’t which is better in general, but what the boosted model buys on your data, and whether that is worth what it costs.
I ran both on the same synthetic readmission data, at two training-set sizes, and scored them on discrimination (AUROC), overall accuracy of the probabilities (Brier score), calibration, and stability. The results give a decision guide you can apply to your own data this afternoon.
The set-up
The data is synthetic, generated with a fixed seed by the script listed at the end: 30,000 discharges with a 30-day readmission rate of 14.4%. I held out 10,000 as a test set and trained on either the first 2,000 or the first 20,000 of the rest. The true risk has structure a plain linear model can’t represent:
- linear effects of age, long-term conditions (LTCs), log length of stay and living in the most deprived quintile;
- prior emergency admissions that raise risk up to four and then stop adding;
- a U-shaped sodium effect, with risk rising for both low and high values;
- an interaction: patients with four or more LTCs who go home without a care package.
Four models, all from scikit-learn:
- Logistic regression: every feature as a straight-line term, standardised, with the default L2 penalty.
- Logistic + terms: the same, plus what an analyst would add after exploring the data: a spline on sodium (
SplineTransformer), prior admissions capped at four, log length of stay, and the LTC × discharge interaction. - Boosting, defaults:
HistGradientBoostingClassifieras it comes. - Boosting, tuned: the same, with a small grid search over learning rate, leaf count and minimum leaf size, scored by log loss with 3-fold cross-validation.
Be clear about one thing. “Logistic + terms” had an unfair advantage: I planted the structure, so I knew which terms to add. On real data you have to find them, and that is where boosting earns its place, as we’ll see.
Discrimination
Data table
| Model | Trained on 2,000 | Trained on 20,000 |
|---|---|---|
| Logistic regression | 0.709 | 0.720 |
| Logistic + terms | 0.789 | 0.791 |
| Boosting, defaults | 0.727 | 0.781 |
| Boosting, tuned | 0.777 | 0.788 |
Synthetic data. Source: projects/article-examples/stats/lr_vs_gb.py.
| Model | AUROC, trained on 2,000 | AUROC, trained on 20,000 |
|---|---|---|
| Logistic regression | 0.709 | 0.720 |
| Logistic + terms | 0.789 | 0.791 |
| Boosting, defaults | 0.727 | 0.781 |
| Boosting, tuned | 0.777 | 0.788 |
Three things stand out.
The plain linear model leaves a lot on the table when the true risk is non-linear. A straight line through a U-shaped sodium effect turns it into a gentle slope, and the interaction is invisible.
Boosting finds most of that structure by itself, given enough data. At 20,000 rows, default boosting reached 0.781 without being told about the sodium curve, the cap on prior admissions or the interaction. At 2,000 rows, the defaults managed only 0.727.
A well-specified logistic regression matches boosting at both sizes. It is also the best model at 2,000 rows. The catch is that you have to know what to specify.
Calibration and the Brier score
AUROC only tells you whether the model ranks patients in the right order. If the output is going to be read as a probability (“this patient has a 30% chance of readmission”), the probabilities have to be right as well. That is calibration, and it is where the small-data boosted model falls apart.
Data table
| Mean predicted risk | Logistic regression | Logistic + terms | Boosting, defaults | Boosting, tuned |
|---|---|---|---|---|
| 0.004 | 4% | |||
| 0.01 | 6% | |||
| 0.018 | 7% | |||
| 0.028 | 8% | |||
| 0.03 | 3% | |||
| 0.041 | 11% | |||
| 0.043 | 4% | |||
| 0.044 | 6% | |||
| 0.045 | 3% | |||
| 0.054 | 4% | |||
| 0.058 | 4% | |||
| 0.06 | 12% | |||
| 0.061 | 6% | |||
| 0.065 | 7% | |||
| 0.068 | 5% | |||
| 0.074 | 7% | |||
| 0.077 | 6% | |||
| 0.08 | 7% | |||
| 0.087 | 8% | |||
| 0.088 | 7% | |||
| 0.09 | 14% | |||
| 0.099 | 9% | 10% | ||
| 0.104 | 10% | |||
| 0.116 | 12% | |||
| 0.121 | 13% | |||
| 0.125 | 13% | |||
| 0.142 | 17% | |||
| 0.144 | 14% | |||
| 0.147 | 20% | |||
| 0.163 | 18% | |||
| 0.174 | 18% | |||
| 0.213 | 26% | |||
| 0.223 | 23% | |||
| 0.235 | 26% | |||
| 0.247 | 24% | |||
| 0.352 | 41% | |||
| 0.424 | 51% | |||
| 0.456 | 55% | |||
| 0.569 | 44% |
Synthetic data. Source: projects/article-examples/stats/lr_vs_gb.py.
| Trained on 2,000 | Brier score | Calibration slope | Mean predicted − observed |
|---|---|---|---|
| Logistic regression | 0.1127 | 1.09 | −0.6 points |
| Logistic + terms | 0.1004 | 1.17 | −1.0 points |
| Boosting, defaults | 0.1135 | 0.53 | −2.4 points |
| Boosting, tuned | 0.1044 | 1.22 | −1.1 points |
The calibration slope comes from regressing the outcome on the logit of the predicted risk. A slope of 1 is ideal. Below 1 means the predictions are too extreme, and above 1 means they are too timid.
Default boosting on 2,000 rows has a slope of 0.53. Its lowest-risk decile averaged a predicted risk of 0.4% against an observed rate of 4.1%, and its highest-risk decile averaged 56.9% predicted against 43.8% observed. It also underestimated the overall rate by 2.4 percentage points. The reason is in the defaults: 100 trees with up to 31 leaves each, and scikit-learn only switches on early stopping automatically when there are more than 10,000 samples. On 2,000 rows the model memorised noise and became overconfident.
Tuning fixed most of that. The grid chose a slower learning rate (0.03) and much smaller trees (7 leaves), and the slope moved to 1.22, slightly timid, as was the regularised logistic model with extra terms at 1.17. At 20,000 rows, every model had a slope between 1.06 and 1.13 and a mean prediction within 0.1 points of the observed rate.
The lesson isn’t that boosting is badly calibrated. It is that boosting’s calibration depends on tuning and data size in a way logistic regression’s mostly doesn’t, so you have to check it every time you retrain.
Stability
A model that gives a patient 18% risk this month and 27% next month, after a routine refresh on slightly different data, is hard to defend. To measure this, I refitted each model on 20 bootstrap resamples of the 2,000-row training set and took the standard deviation of each test patient’s predicted risk across the 20 fits.
| Trained on 2,000 | Median SD of a patient’s predicted risk |
|---|---|
| Logistic regression | 1.9 points |
| Logistic + terms | 1.7 points |
| Boosting, defaults | 4.4 points |
Resampling the same data moved a typical patient’s boosted risk by more than twice as much as their logistic risk. Tuning and more data narrow the gap, but trees split on thresholds, and small changes to the data move the thresholds.
Explanation
The plain logistic model’s coefficients are its explanation. On the 2,000-row fit, each standard deviation of prior admissions multiplied the odds of readmission by 1.51, of LTCs by 1.32, and of age by 1.26. A clinician can argue with those numbers, and should.
Boosting has no coefficients. You explain it after the fact, with permutation importance (sklearn.inspection.permutation_importance), partial dependence (sklearn.inspection.partial_dependence), or SHAP-style attributions. These are useful, but they are summaries of a model you can’t read directly, and they need care when features are correlated. Explaining models to stakeholders works through them on a readmission model, correlated features included.
The explanation tools also offer a route back to a simpler model. If boosting clearly beats a plain logistic regression, plot partial dependence for the top features. A U-shape or a plateau tells you which term to add. Add it, refit the logistic model, and see how much of the gap closes. Here, adding four terms closed all of it.
A decision guide
Start with logistic regression when:
- you have a few thousand rows or fewer, or few outcome events;
- people need to see and challenge the effect of each factor;
- the score has to run somewhere simple: SQL, DAX, a spreadsheet, a paper form;
- scores need to stay stable between refreshes.
Consider gradient boosting when:
- you have tens of thousands of rows or more;
- it beats a well-specified logistic model (not just a plain one) by a margin that changes decisions;
- you can monitor calibration in production and recalibrate or retrain when it drifts;
- you have time for tuning and for explaining a model that has no coefficients.
Whichever you choose:
- Fit both. Treat the boosted model as a diagnostic even if you don’t ship it: a large gap means structure the linear model misses.
- Report calibration alongside AUROC, on held-out data, every time.
- Check stability with a bootstrap or two refits before a model goes near a patient list.
- Compare against the well-specified linear model, not the naive one. Beating a straw man proves nothing.
The readmission risk model project applies these checks to a larger synthetic cohort with the messier problems real data brings: missing values, a trend over time and a leakage trap.
Reproduce
All the numbers above come from projects/article-examples/stats/lr_vs_gb.py, which simulates the data, fits the four models at both sizes, runs the bootstrap and writes both charts. It takes about half a minute:
uv run python projects/article-examples/stats/lr_vs_gb.py
See it in a project
Tags
- logistic-regression
- gradient-boosting
- scikit-learn
- calibration
- model-selection