Referral triage text classifier
A TF-IDF and logistic regression baseline that routes synthetic GP referral letters by specialty and urgency, with a human-review threshold, calibration checks and error analysis. It's the bar an LLM would have to beat.
Synthetic data Every record here is generated. No real patient or organisational data is used.
- Urgency macro-F1 on unseen practices (keyword rules: 0.670)
- 0.944
- Urgent and suspected-cancer letters sent to a person (471 of 483)
- 97.5%
- Letters that could skip the first triage read
- 55.9%
- Missed fast-track letters whose label I had moved on purpose
- 9 of 12
Every referral letter that reaches a hospital has to be read by someone who decides which specialty should see the patient, and how soon. Most letters are routine. A few describe a possible cancer, and those can’t wait in the same queue.
This project builds the classical baseline for that job, TF-IDF and logistic regression, compares it with keyword rules, and evaluates it the way a triage tool would need to be: recall on the urgent classes, calibration, a threshold for which letters a person must read, and the letters it gets wrong. Any LLM proposed for this job would have to beat it.
Everything here is synthetic. The letters are generated from templates by the project code with a fixed seed. No real patient, clinician or practice appears in them. The labels follow simple rules I wrote, loosely modelled on UK referral practice. They are not clinical guidance.
In short: on 11 practices the model never saw, it got the urgency right on 95.6% of letters, against 66.4% for keyword rules. With a threshold chosen on validation practices, 55.9% of letters could skip the first triage read while 471 of 483 urgent or suspected-cancer letters still went to a person.
The problem
A referral hub, or a consultant vetting their list, gives each letter a specialty and a priority: routine, urgent, or suspected cancer (once called two-week wait, now measured in England against the 28-day Faster Diagnosis Standard). A missed suspected-cancer letter delays a diagnosis; a routine letter marked urgent takes someone else’s slot. Both cost something, but not the same.
So the model’s job is narrow on purpose. It can only say “this looks routine, and it belongs to this specialty”. Anything that might be urgent or cancer goes to a person. I built the data generator, the models and the evaluation.
What’s in the data
6,000 letters from 40 fictional practices, with a median of 43 words (30 to 58 for the middle 80%). There are eight specialties, from orthopaedics (20.2% of letters) to neurology (6.4%), and three urgencies: routine 61.3%, suspected cancer 20.4%, urgent 18.2%.
Each letter comes from one of 50 presentations, such as knee osteoarthritis, a changing mole or haemoptysis. On top of that I planted what makes real letters hard, and recorded each one in a column so errors can be traced back to it:
| Difficulty | What I planted | Letters |
|---|---|---|
| Negation | “Denies wt loss, haematuria or fever”; “Does not meet 2WW criteria” | 50.4% |
| Red flags | A red flag raises the urgency (“ongoing reflux and weight loss”); letters without it often deny it | 5.4% |
| Age thresholds | Change in bowel habit: suspected cancer from 60. Visible haematuria: suspected cancer from 45, urgent below | 4.3% |
| Typos | A typo rate per practice, from none to 6% of longer words | 37.9% |
| Abbreviations | SOB, PMH, c/o, wt loss, 3/52, at a rate set per practice | varies |
| Other history | Past history and medication from another specialty | 50.0% |
| Rare phrasing | “Dysphonia unresolved after a month” for “persistent hoarseness” | 3.0% |
| American spelling | “Hematuria”, “esophageal”, from 12 of the 40 practices | 2.7% |
| Label noise | 1.5% of urgency and 1% of specialty labels moved, standing in for triager disagreement | 2.8% |
A typical letter (L00043 in the download):
Dear Team,
Thank you for seeing this 70 year old man with lower urinary tract symptoms for 2 weeks.
Prostate feels smooth and benign. PSA within the normal range. Associated urgency.
Denies wt loss, haematuria or fever. Possible BPH. Past medical history: gastro-oesophageal
reflux. Please see and advise on further management.
Best wishes.
Template letters are much cleaner than real ones. Treat the absolute scores as optimistic, and the comparisons and failure patterns as the useful part.
Approach
Split by practice. Whole practices go to one split: 21 practices (3,557 letters) for training, 8 (1,156) for validation and 11 (1,287) for test. Each practice has its own greeting, abbreviations and typo rate, so a random split would let the model meet a practice’s style again at test. Every choice, including C, the feature set, class weights and the review threshold, was made on the validation practices. The test practices were only used to report.
Keyword rules. The kind an analyst writes in an afternoon: “palpitations” or “murmur” means cardiology; “2ww”, “weight loss”, “haematuria” or “lump” means suspected cancer. I didn’t edit them after the first run.
The model. One scikit-learn pipeline per label: TF-IDF on word 1-2 grams next to TF-IDF on character 2-5 grams within word boundaries, into a multinomial logistic regression.
def features(
word: tuple[int, int] | None, char: tuple[int, int] | None, stop_words: str | None = None
) -> FeatureUnion:
"""TF-IDF on word n-grams, character n-grams within word boundaries, or both side by side."""
parts = []
if word:
parts.append(
(
"word",
TfidfVectorizer(
ngram_range=word, min_df=2, sublinear_tf=True, stop_words=stop_words
),
)
)
if char:
parts.append(
(
"char",
TfidfVectorizer(analyzer="char_wb", ngram_range=char, min_df=2, sublinear_tf=True),
)
)
return FeatureUnion(parts)
Of five feature sets, words and characters together had the lowest validation log loss for both labels, without class weights: C = 30 for specialty and C = 10 for urgency. Build a text classification baseline before you reach for an LLM compares the feature sets.
Results
| Test practices | Specialty macro-F1 | Specialty accuracy | Urgency macro-F1 | Urgency accuracy |
|---|---|---|---|---|
| Keyword rules | 0.810 | 81.6% | 0.670 | 66.4% |
| TF-IDF + logistic regression | 0.987 | 98.8% | 0.944 | 95.6% |
Data table
| Class | TF-IDF + logistic regression | Keyword rules |
|---|---|---|
| Orthopaedics | 0.99 | 0.81 |
| Dermatology | 0.99 | 0.82 |
| Gastroenterology | 0.99 | 0.86 |
| Cardiology | 0.99 | 0.81 |
| Urology | 0.98 | 0.88 |
| ENT | 0.98 | 0.79 |
| Respiratory | 0.98 | 0.74 |
| Neurology | 1.00 | 0.76 |
| Class | TF-IDF + logistic regression | Keyword rules |
|---|---|---|
| Routine | 0.97 | 0.71 |
| Urgent | 0.93 | 0.78 |
| Suspected cancer | 0.93 | 0.52 |
Synthetic referral letters. Source: projects/referral-triage-classifier.
Specialty is the easy half: the model’s lowest F1 is ENT at 0.977. Switch the heatmap to the keyword rules and their most common mistake appears in the first column. Orthopaedics is where a letter lands when no keyword matches or two specialties tie.
Data table
| True specialty | Ortho | Derm | Gastro | Cardio | Uro | ENT | Resp | Neuro |
|---|---|---|---|---|---|---|---|---|
| Orthopaedics | 98% | 1% | 0% | 0% | 0% | 0% | 0% | 0% |
| Dermatology | 0% | 100% | 0% | 0% | 0% | 0% | 0% | 0% |
| Gastroenterology | 0% | 0% | 97% | 0% | 1% | 1% | 0% | 0% |
| Cardiology | 0% | 0% | 0% | 99% | 0% | 1% | 0% | 0% |
| Urology | 0% | 0% | 0% | 0% | 100% | 0% | 0% | 0% |
| ENT | 0% | 0% | 0% | 0% | 0% | 99% | 1% | 0% |
| Respiratory | 0% | 0% | 0% | 2% | 1% | 0% | 97% | 0% |
| Neurology | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 100% |
| True specialty | Ortho | Derm | Gastro | Cardio | Uro | ENT | Resp | Neuro |
|---|---|---|---|---|---|---|---|---|
| Orthopaedics | 93% | 2% | 1% | 0% | 0% | 0% | 1% | 4% |
| Dermatology | 12% | 86% | 0% | 0% | 0% | 0% | 2% | 0% |
| Gastroenterology | 7% | 6% | 84% | 1% | 1% | 1% | 1% | 0% |
| Cardiology | 8% | 7% | 4% | 74% | 0% | 1% | 7% | 0% |
| Urology | 15% | 2% | 2% | 0% | 81% | 0% | 0% | 0% |
| ENT | 7% | 6% | 6% | 0% | 0% | 77% | 1% | 4% |
| Respiratory | 5% | 5% | 4% | 8% | 0% | 11% | 68% | 0% |
| Neurology | 6% | 12% | 4% | 0% | 0% | 4% | 0% | 74% |
Synthetic referral letters. Source: projects/referral-triage-classifier. Rows sum to 100%. Columns use short names.
Urgency is where the rules fall apart. They called 301 of the 804 routine letters suspected cancer, because their cancer words turn up in routine letters too: “PSA within the normal range”, “Denies wt loss, haematuria or fever”, “previous breast cancer, in remission”.
| Model, test practices | Precision | Recall | Letters |
|---|---|---|---|
| Routine | 0.951 | 0.993 | 804 |
| Urgent | 0.982 | 0.888 | 240 |
| Suspected cancer | 0.952 | 0.905 | 243 |
A suspected-cancer recall of 0.905 means 21 of those letters, plus 20 urgent ones, were predicted routine. Taking the most likely class is the wrong way to use this model.
A threshold for the human read
Instead, each letter gets a probability of being urgent or suspected cancer, and anything at or above a threshold goes to a person. On the validation practices I chose the highest threshold that still sent at least 99% of suspected-cancer letters and 95% of urgent letters to a human. Those targets are my assumptions about relative cost. A real service would set them with its clinical safety officer.
The rule picked 0.11. On the test practices:
- 719 of 1,287 letters (55.9%) were auto-routed as routine;
- 471 of 483 urgent or suspected-cancer letters (97.5%; 95% Wilson interval 95.7% to 98.6%) went to a person, including 242 of 243 suspected-cancer letters;
- 98.9% of the auto-routed letters had the right specialty.
Data table
| Share of letters sent to human triage | Suspected cancer | Urgent |
|---|---|---|
| 0.198 | 51% | 55% |
| 0.228 | 57% | 64% |
| 0.246 | 62% | 68% |
| 0.255 | 65% | 70% |
| 0.261 | 67% | 72% |
| 0.272 | 70% | 75% |
| 0.276 | 71% | 76% |
| 0.283 | 73% | 78% |
| 0.287 | 74% | 79% |
| 0.291 | 74% | 81% |
| 0.294 | 75% | 82% |
| 0.298 | 76% | 83% |
| 0.303 | 77% | 84% |
| 0.307 | 78% | 85% |
| 0.311 | 79% | 86% |
| 0.313 | 79% | 87% |
| 0.315 | 79% | 88% |
| 0.316 | 80% | 88% |
| 0.317 | 80% | 88% |
| 0.319 | 82% | 88% |
| 0.322 | 83% | 88% |
| 0.323 | 83% | 88% |
| 0.326 | 84% | 89% |
| 0.327 | 84% | 89% |
| 0.329 | 85% | 89% |
| 0.331 | 86% | 89% |
| 0.333 | 86% | 90% |
| 0.334 | 86% | 90% |
| 0.335 | 86% | 90% |
| 0.336 | 87% | 91% |
| 0.337 | 87% | 91% |
| 0.341 | 88% | 91% |
| 0.343 | 90% | 91% |
| 0.347 | 91% | 92% |
| 0.348 | 91% | 92% |
| 0.349 | 92% | 92% |
| 0.35 | 92% | 92% |
| 0.351 | 92% | 92% |
| 0.353 | 93% | 92% |
| 0.355 | 93% | 92% |
| 0.358 | 93% | 93% |
| 0.36 | 93% | 93% |
| 0.361 | 94% | 93% |
| 0.363 | 95% | 93% |
| 0.367 | 95% | 94% |
| 0.368 | 96% | 94% |
| 0.371 | 96% | 94% |
| 0.375 | 97% | 94% |
| 0.377 | 97% | 94% |
| 0.378 | 97% | 94% |
| 0.381 | 97% | 94% |
| 0.382 | 97% | 94% |
| 0.384 | 97% | 94% |
| 0.387 | 98% | 94% |
| 0.39 | 98% | 94% |
| 0.395 | 98% | 95% |
| 0.399 | 98% | 95% |
| 0.402 | 98% | 95% |
| 0.403 | 98% | 95% |
| 0.406 | 99% | 95% |
| 0.411 | 99% | 95% |
| 0.414 | 99% | 95% |
| 0.418 | 99% | 95% |
| 0.421 | 99% | 95% |
| 0.428 | 99% | 95% |
| 0.436 | 100% | 95% |
| 0.441 | 100% | 95% |
| 0.451 | 100% | 96% |
| 0.465 | 100% | 96% |
| 0.482 | 100% | 97% |
| 0.497 | 100% | 97% |
| 0.521 | 100% | 97% |
| 0.548 | 100% | 98% |
| 0.58 | 100% | 98% |
| 0.624 | 100% | 98% |
| 0.698 | 100% | 98% |
| 0.793 | 100% | 99% |
Synthetic referral letters. Source: projects/referral-triage-classifier. Threshold chosen on the validation practices, shown on test.
This curve is the conversation to have with a service. At a threshold of 0.5, only 34.9% of letters go to a person, but 8% of suspected-cancer letters are missed. At 0.05, 54.8% go to a person and no suspected-cancer letter in this test set is missed. Each step towards safety costs reading time, and the chart shows how much.
Of the 12 fast-track letters the rule let through, 9 had labels I had moved on purpose. Eight were routine letters moved up to urgent, like a sleep apnoea referral with no red flags. Real data has no label_changed column, so those eight would look like model failures. That’s why a test set’s labels need checking, as evaluating LLM features with a test set explains.
The other four were red flags the model underweighted: palpitations “plus CP”, “There has been some chest pain” after palpitations, “recurrent migraines, with focal neurology”, and reflux with “some dysphagia”. That last letter is the ninth moved label: suspected cancer moved down to urgent, so the model missed it under either label.
Calibration
The threshold was chosen by recall, so it works even if the probabilities are off. Calibration matters because people will read a score of 0.11 as an 11% chance.
Data table
| Predicted probability | No class weights | Balanced class weights |
|---|---|---|
| 0.001 | 2% | |
| 0.003 | 1% | |
| 0.004 | 1% | |
| 0.007 | 2% | |
| 0.01 | 1% | |
| 0.014 | 1% | |
| 0.022 | 1% | |
| 0.027 | 0% | |
| 0.047 | 2% | |
| 0.05 | 3% | |
| 0.112 | 7% | |
| 0.129 | 8% | |
| 0.52 | 63% | |
| 0.643 | 62% | |
| 0.953 | 100% | |
| 0.987 | 100% | |
| 0.995 | 99% | |
| 0.999 | 99% | |
| 1 | 100% | 100% |
Synthetic referral letters. Source: projects/referral-triage-classifier.
Without class weights, the mean predicted probability of urgent or cancer was 36.8% against 37.5% observed (Brier score 0.0303). The weak spot is the middle: the group with a mean prediction of 52% was 63% urgent or cancer. Balanced weights did slightly better on test (Brier 0.0290), but validation preferred no weights, and switching after seeing test results would be tuning on test. With a rarer outcome, weights do real damage, as calibration before AUC shows.
Where it fails
Because every difficulty was planted, errors can be grouped by cause.
Data table
| Letters with (count) | TF-IDF + logistic regression | Keyword rules |
|---|---|---|
| Negated red flag or reassurance (640) | 3% | 52% |
| Red flag raises the urgency (73) | 32% | 15% |
| Rare phrasing of the complaint (53) | 8% | 34% |
| At least one typo (270) | 5% | 37% |
| Urgency depends on age (61) | 18% | 39% |
| History from another specialty (660) | 4% | 32% |
| American spelling (65) | 9% | 43% |
| Urgency or specialty label moved on purpose (27) | 63% | 74% |
| All test letters (1287) | 4% | 34% |
Synthetic referral letters. Source: projects/referral-triage-classifier. A letter can have several difficulties.
- Negation is mostly handled. On the 640 letters with negation, the model’s urgency error was 3.4%, against 51.7% for the rules.
- Red flags are its worst case. Where a red flag raised the urgency, the model was wrong on 31.5% of 73 letters, and the keyword rules on 15.1%. In the training letters, 49 of the 54 sentences mentioning “focal neurology” put a negation word before it, and 229 of 256 for “weight loss”, so the model learnt them as routine words. A 29-year-old with “ongoing reflux and weight loss” got a routine probability of 0.744. The threshold still sent it to a person.
- Age thresholds. On the 61 age-dependent letters, the error was 18.0%. A 54-year-old with a change in bowel habit, routine in these labels, got a suspected-cancer probability of 0.811. TF-IDF sees “54” as a token, not a number.
- Specialty errors are mostly label noise. Of 16 specialty errors on test, 10 were letters whose specialty label I had moved. On the validation practices, all 14 errors among letters with a top probability of 0.9 or more were moved labels. A confidence band wouldn’t catch them, because the letter clearly says something other than its label.
Data table
| Word or word pair | Coefficient |
|---|---|
| loss for | 2.4 |
| 2ww | 2.0 |
| 2ww referral | 1.9 |
| cancer | 1.9 |
| with | 1.7 |
| lump | 1.6 |
| has had | 1.5 |
| reports | 1.5 |
| in bowel | 1.5 |
| suspected cancer | 1.4 |
| Word or word pair | Coefficient |
|---|---|
| who had | 1.9 |
| see urgently | 1.8 |
| urgently | 1.8 |
| prompt review | 1.8 |
| prompt | 1.8 |
| appreciate prompt | 1.7 |
| flare | 1.7 |
| an urgent | 1.4 |
| urgent appointment | 1.4 |
| pain 52 | 1.3 |
| Word or word pair | Coefficient |
|---|---|
| or | 2.8 |
| no | 2.2 |
| loss or | 1.9 |
| denies | 1.7 |
| nil | 1.7 |
| utis | 1.5 |
| oa | 1.4 |
| or dysphagia | 1.4 |
| left hip | 1.3 |
| knee | 1.2 |
Synthetic referral letters. Source: projects/referral-triage-classifier. Character n-grams also carry weight but are not shown.
The coefficients show what the model leans on. “No”, “or” and “denies” push towards routine because negation lists live mostly in routine letters. “2WW referral” pushes towards suspected cancer, as it should, but so do “has had” and “reports”, because the generator tends to introduce red flags that way. Real letters have their own quirks, like a practice template, and reading the coefficients is the cheapest way to find them.
Where an LLM fits
Several failure clusters are the kind a language model might handle better: an age against a threshold, a negation covering a list, “dysphonia” meaning hoarseness, a red flag written like an ordinary symptom. I haven’t run one here, and nothing on this page is an LLM result. I’d run the comparison like this:
- Same test practices, labels and metrics, including fast-track recall at the operating point and the share of letters that skip the human read.
- A score, not just a label, to set the threshold on: token log-probabilities if the API exposes them, or a stated confidence that has been checked for calibration on the validation practices.
- Prompt and threshold chosen on validation, then one run on test, and a rerun on every prompt or model version change.
- The full cost: latency, price per letter, and whether answers change between runs.
Governance comes first. Referral letters are special category health data under UK GDPR, so an external API needs a lawful basis, a data protection impact assessment, a processing agreement covering retention and training use, and the Caldicott Guardian involved. In England, clinical decision support falls under the NHS clinical risk management standards DCB0129 and DCB0160, and depending on its intended purpose, triage software can qualify as a medical device under MHRA guidance.
That’s why the baseline is the bar. It runs on a laptop, no data leaves the organisation, the same letter always gets the same answer, and its weights can be read. An LLM has to beat 97.5% fast-track recall at 55.9% auto-routing, and be worth the governance work. If it is, Content gate: LLM review with code-owned verdicts shows the pattern I’d use: the model answers typed questions, and plain code decides.
Limits
- Template text is easier than real text. 98.8% specialty accuracy won’t survive real letters. The ranking of methods and the failure patterns are what transfer.
- The labels are rules I wrote, simplified from UK practice. They’re not clinical guidance.
- Not fit for real triage. This is a benchmark on synthetic letters. Any real use would need evaluation on governed real letters, a clinical safety case under DCB0129 and DCB0160, and the service’s own decision on which letters, if any, can skip a clinician’s read.
- One test split. 243 suspected-cancer letters give a recall interval about 2 points wide, and the 65 American-spelling and 53 rare-phrasing letters move a lot with a few errors.
What I’d do next
- Features for the known gaps: an extracted age with the two thresholds, and a negation-scope rule that marks terms after “no”, “denies” or “nil” up to the end of the sentence.
- Repeated grouped splits, to put intervals on the comparison between feature sets.
- The LLM comparison above, with governance in place first, if the gaps justify it.
- Monitoring after launch: the auto-routed share and how often triagers overturn the model, by practice, since a new practice template would shift both.
Run it yourself
uv run python projects/referral-triage-classifier/run.py
It regenerates the letters, fits every model, chooses the threshold, rewrites every chart and the download on this page, and prints the tables quoted here, in about three to four minutes. A rerun gives byte-identical files. The code is in the download below, with a README that explains each file.
Built with
- Python
- scikit-learn
- pandas
- statsmodels
- NumPy
Downloads
-
Test-practice letters with predictions
1,287 synthetic letters from 11 practices: labels, planted-difficulty flags, keyword and model predictions, probabilities and the routing decision.
- Project code