-
Explaining models to stakeholders
How to explain a readmission risk model to clinicians and managers. Global and local explanations, permutation importance and partial dependence in scikit-learn, why correlated features mislead, what-if examples, uncertainty, and model cards.
-
Data leakage in healthcare ML, and how to catch it
Where leakage hides in health data, from post-discharge fields to time travel in lag features, how to detect it, and two worked examples from a readmission model.
-
Calibration before AUC
Why a risk score used for decisions needs probabilities that match reality, and how to check and fix calibration in scikit-learn, with numbers from a readmission model.
-
Logistic regression or gradient boosting for tabular health data
A head-to-head on synthetic readmission data with scikit-learn, comparing AUROC, Brier score, calibration and stability, and a decision guide for choosing between them.
-
Build a text classification baseline before you reach for an LLM
Why TF-IDF and logistic regression should come first for classifying clinical or operational text, how to handle negation, abbreviations and typos, and how to tell when an LLM is likely to beat it.