Build a text classification baseline before you reach for an LLM
Why TF-IDF and logistic regression should come first for classifying clinical or operational text, how to handle negation, abbreviations and typos, and how to tell when an LLM is likely to beat it.
If you have a pile of labelled text and a question like “which team should get this?” or “how urgent is it?”, the first model to build is TF-IDF with logistic regression. It takes an afternoon, runs on a laptop, gives the same answer every time, and its coefficients tell you what it’s using. Most usefully, it gives you two things any LLM proposal needs: the number to beat, and a map of where the problem is actually hard.
The examples here come from my referral triage text classifier, which classifies 6,000 synthetic GP referral letters by specialty and urgency. The letters are generated from templates, with negation, abbreviations, typos and label noise planted on purpose. No real patient data is involved, and the labels are not clinical guidance. All the numbers are on 1,287 letters from 11 practices the models never saw.
What a baseline is good at
A bag-of-words model is strong wherever the label shows up in the vocabulary. “Palpitations” means cardiology, “tinnitus” means ENT, and “2WW referral” means suspected cancer. On specialty routing, the baseline reached 98.8% accuracy, against 81.6% for a set of keyword rules. On urgency it reached 95.6%, against 66.4%.
It also brings things an LLM doesn’t give you for free:
- Probabilities you can check. Logistic regression outputs a probability for each class, and you can test whether they’re calibrated and set a threshold on them.
- Speed and cost. Fitting takes seconds and scoring is effectively free, so you can retrain it every week or run a hundred experiments.
- Determinism. The same letter gets the same answer, which makes regression testing simple.
- Inspectable weights. You can list the words that push towards each class and spot a shortcut before it reaches production.
- No data leaves the building. For clinical text, that matters more than anything above.
Build it properly
Here are the features from the project. Word n-grams and character n-grams sit side by side in a FeatureUnion:
def features(
word: tuple[int, int] | None, char: tuple[int, int] | None, stop_words: str | None = None
) -> FeatureUnion:
"""TF-IDF on word n-grams, character n-grams within word boundaries, or both side by side."""
parts = []
if word:
parts.append(
(
"word",
TfidfVectorizer(
ngram_range=word, min_df=2, sublinear_tf=True, stop_words=stop_words
),
)
)
if char:
parts.append(
(
"char",
TfidfVectorizer(analyzer="char_wb", ngram_range=char, min_df=2, sublinear_tf=True),
)
)
return FeatureUnion(parts)
sublinear_tf=True replaces a term’s count with 1 + log(count), so a word repeated five times doesn’t count five times as much. min_df=2 drops n-grams that appear in only one training letter. analyzer="char_wb" builds character n-grams only from text inside word boundaries, padding each word with spaces, so “?melanoma” produces " ?m", “?me” and “mel”.
Two choices matter more than the feature settings.
Split by source, not by row. Each practice in the project has its own greeting, abbreviation habits and typo rate. A random split would let the model learn a practice’s style from training and meet it again at test. Split by whichever grouping your text comes from (practice, team, author, template) or by time. Data leakage in healthcare ML, and how to catch it applies the same logic more generally.
Tune on validation, report on test. I chose C, the feature set and class weights by log loss on 8 validation practices, and only then scored the 11 test practices once.
Negation and abbreviations
Clinical text is full of pertinent negatives: “No weight loss”, “Denies haematuria”, “Nil focal neurology”. A keyword rule sees “weight loss” and fires. In the project, keyword rules called 301 of 804 routine letters suspected cancer, and got the urgency wrong on 51.7% of the letters with negation. The TF-IDF model got 3.4% of them wrong. “No”, “denies” and pairs like “loss or” are among its largest coefficients towards routine, so the negation travels with the symptom.
That protection is easy to throw away. Don’t remove stop words. scikit-learn’s built-in English list includes “no”, “not”, “nor”, “without”, “never”, “none” and “cannot”. Remove them and “no weight loss” becomes “weight loss”.
Data table
| Features | Urgency | Specialty |
|---|---|---|
| Keyword rules | 33.6% | 18.4% |
| Word unigrams | 4.4% | 1.3% |
| Word 1-2 grams, stop words removed | 5.7% | 1.2% |
| Word 1-2 grams | 5.0% | 1.6% |
| Character 2-5 grams | 4.2% | 1.5% |
| Word and character n-grams | 4.4% | 1.2% |
| Features | Urgency | Specialty |
|---|---|---|
| Keyword rules | 51.7% | 14.7% |
| Word unigrams | 3.6% | 0.9% |
| Word 1-2 grams, stop words removed | 4.2% | 0.8% |
| Word 1-2 grams | 3.7% | 0.8% |
| Character 2-5 grams | 3.9% | 1.3% |
| Word and character n-grams | 3.4% | 0.6% |
| Features | Urgency | Specialty |
|---|---|---|
| Keyword rules | 15.1% | 16.4% |
| Word unigrams | 30.1% | 0.0% |
| Word 1-2 grams, stop words removed | 47.9% | 0.0% |
| Word 1-2 grams | 38.4% | 0.0% |
| Character 2-5 grams | 23.3% | 0.0% |
| Word and character n-grams | 31.5% | 0.0% |
| Features | Urgency | Specialty |
|---|---|---|
| Keyword rules | 36.7% | 16.7% |
| Word unigrams | 5.6% | 1.5% |
| Word 1-2 grams, stop words removed | 6.3% | 1.1% |
| Word 1-2 grams | 6.3% | 1.1% |
| Character 2-5 grams | 4.4% | 1.9% |
| Word and character n-grams | 5.2% | 1.1% |
| Features | Urgency | Specialty |
|---|---|---|
| Keyword rules | 43.1% | 23.1% |
| Word unigrams | 9.2% | 3.1% |
| Word 1-2 grams, stop words removed | 10.8% | 4.6% |
| Word 1-2 grams | 9.2% | 4.6% |
| Character 2-5 grams | 6.2% | 1.5% |
| Word and character n-grams | 9.2% | 3.1% |
Synthetic referral letters. Source: projects/referral-triage-classifier. No class weights; C chosen on the validation practices.
With stop words removed, urgency error on the test practices rose from 5.0% to 5.7% for the same word 1-2 gram model. On the 73 letters where a red flag was present and raised the urgency, it rose from 38.4% to 47.9%. With “no” gone, a present red flag and an absent one can look identical.
Bigrams only reach one word. “No weight loss, night sweats or haemoptysis” negates three things, and “haemoptysis” sits six words after the “no”. For that, either add a simple negation-scope feature (mark terms between “no”, “denies” or “nil” and the end of the sentence) or accept it as a known gap.
For abbreviations, the data usually teaches the model that “SOB” and “shortness of breath” mean the same, as long as both appear in training. What breaks is an abbreviation or spelling the training set never saw, and that’s where character n-grams come in.
Character n-grams for typos and spelling
Word features treat “haemoptysis”, “hemoptysis” and “haemoptyiss” as three unrelated tokens. Character n-grams share most of their pieces. On the 270 test letters with at least one typo, character 2-5 grams alone had an urgency error of 4.4%, against 6.3% for word 1-2 grams. On the 65 letters with American spellings from practices the model hadn’t seen, it was 6.2% against 9.2%.
But be honest about the size of the effect. Overall, every model with word or character features landed between 4.2% and 5.7% urgency error. The model chosen on validation, words and characters together, matched plain word unigrams on test at 4.4%. Each got 10 letters right that the other got wrong (McNemar’s exact test, p = 1.0). Character n-grams are cheap insurance against messy text, not a guaranteed gain. Measure them on your own data.
Class imbalance
The urgency split in the project is 61% routine, 18% urgent and 20% suspected cancer. Three habits matter more than any resampling trick:
- Report macro-F1 and per-class recall, not accuracy. Accuracy is dominated by the big class.
- Choose class weights on validation, like any other setting.
class_weight="balanced"weights each class inversely to its frequency. Here validation log loss preferred no weights, and the difference on test was small. With a rarer outcome, balanced weights inflate the probabilities of the rare class, which matters if anyone reads them as risks: see calibration before AUC. - Use a threshold, not the most likely class. Taking the most likely class left 41 urgent or suspected-cancer letters predicted routine. Sending every letter with an urgent-or-cancer probability of 0.11 or more to a person cut that to 12, while still letting 55.9% of letters skip the first read.
Error analysis
A single score hides the part you need. Group the errors by cause, then read them.
In the project, the causes were planted, so the grouping is exact. With real data, tag a sample of errors by hand: negation, abbreviation, spelling, missing information, label disagreement. A few patterns showed up:
- Red flags that look like ordinary symptoms. “Recurrent migraines, with focal neurology” got a routine probability high enough to skip review. In the training letters, “focal neurology” mostly appears after “no”, so the model learnt it as a routine phrase.
- Numbers. A change in bowel habit is routine under 60 and suspected cancer from 60 in these labels. The model’s urgency error on age-dependent letters was 18.0%, against 4.4% overall. TF-IDF sees “54” as a token, not a number.
- Label noise. Of the 12 urgent or cancer letters the threshold let through, 9 had labels I had moved on purpose, and 8 of those were routine letters moved up to urgent. With real data those eight would look like model failures. Read a sample of “errors” before you believe them.
Then read the coefficients. The project’s top_terms lists the word n-grams with the largest weights for a class, limited to those in at least 1% of training letters:
words = coefs[coefs.index.str.startswith("word__")]
words.index = words.index.str.removeprefix("word__")
common = df[df >= MIN_TERM_SHARE * len(texts)].index
return words[words.index.isin(common)].sort_values(ascending=False).head(n)
For suspected cancer, “2WW referral” and “lump” rank high, as they should. So do “has had” and “reports”, because the letter generator tends to introduce red flags that way. In real text, the equivalent is a practice template or a phrase one clinician always uses. Coefficients are the cheapest shortcut detector you have.
When an LLM is likely to beat it
The baseline’s failures point to where a language model has a real chance:
- Reasoning over numbers and conditions, such as an age against a threshold, or a result against a reference range.
- Negation and scope across a sentence: “No visual disturbance but reports weakness”.
- Synonyms and phrasing never seen in training: “dysphonia” for hoarseness, “velcro crackles”.
- Long documents with several problems, where the label depends on which one matters.
- Few labelled examples, or a label set that changes too often to keep retraining.
And where it probably won’t: tasks driven by vocabulary with plenty of labelled data, like specialty routing here, and cases where the baseline’s remaining errors are mostly label noise. No model beats the labels.
Whatever you test, test it on the same held-out set with the same metrics, including recall at the operating point and the workload it saves. Evaluating LLM features with a test set covers how to build that set, and for clinical text, the governance has to be in place before any letter goes near an external API.
A checklist
- Split by source or time, and keep a validation set for every choice.
- TF-IDF with word 1-2 grams,
sublinear_tf=True, no stop-word removal. Addchar_wbn-grams if the text is messy. - Logistic regression, with C and class weights chosen on validation log loss.
- Report macro-F1, per-class precision and recall, and a confusion matrix, next to a keyword-rule baseline.
- Check calibration, then set thresholds on validation according to the cost of each error.
- Group errors by cause, read them, and read the coefficients.
- Write down what the baseline can’t do. That list is the case for, or against, an LLM.
See it in a project
Tags
- text-classification
- tf-idf
- logistic-regression
- scikit-learn
- llm
- error-analysis