Behnam Analytics

Writing AI workflows & prompting

Evaluating LLM features with a test set

How to build a labelled evaluation set before shipping an LLM feature, choose metrics that match the risk, catch regressions when prompts or models change, and keep clinical text governed throughout.

Behnam Ebrahimi 8 min read

An LLM feature without a labelled test set is a demo. It works on the ten examples someone tried, and nobody can say how often it’s wrong, on which cases, or whether last week’s prompt change made it better or worse. Build the test set before the prompt, and most of the hard questions about the feature get answered on the way.

This applies to any feature that turns text into a decision: routing a referral, tagging a complaint, extracting a date from a letter, flagging a record for review. I’ll use my referral triage text classifier for the numbers. It’s a classical model on synthetic letters, not an LLM, but the evaluation is the same, and that’s the point.

Start from the decision, not the model

Before sampling anything, write down three things:

  1. What decision the output drives. “Letters the model calls routine skip the first triage read” is a decision. “Classify letters by urgency” isn’t one yet.
  2. Which error costs more. A missed suspected-cancer referral is a delay in diagnosis. A routine letter sent for review costs a few minutes of a clinician’s time. Those errors shouldn’t be weighted the same.
  3. What happens to the output next. Does a person see it? Can they overrule it? Is there a fallback when the model fails or times out?

Those answers decide which metrics matter, how many examples you need, and what “good enough” means.

Sample the set from real inputs

The test set should look like the traffic the feature will see, not like the examples that were easy to find.

  • Sample from the actual source, across the period and the sources the feature will serve. If letters come from 40 practices, the set should include many of them, and ideally whole practices kept out of any development work, as in the project’s split.
  • Stratify where it matters. Rare but costly cases need enough examples to measure. It’s fine to oversample them, as long as you report recall per class and reweight precision and any pooled figure to the natural mix. Precision depends on prevalence, so an oversampled class looks more precise than it will be in live use.
  • Keep a separate development set. Prompt iteration, few-shot examples and threshold choices all happen on the development set. The test set gets used rarely, ideally once per candidate you’re seriously considering.

How big? Big enough that the number you care about has a useful interval. For recall, what matters is the number of positive cases, not the size of the whole set.

How sure is a measured recall of 90%?95% Wilson interval by the number of positive cases in the test set
Data table
Positive cases in the test setMeasured recallMeasured recall (low)Measured recall (high)
1090%60%98%
2090%70%97%
3090%74%96%
4090%77%96%
5090%79%96%
7590%82%95%
10090%83%94%
15090%84%94%
20090%85%93%
30090%86%93%
40090%87%93%
50090%87%92%
75090%88%92%
100090%88%92%

Calculated with statsmodels proportion_confint(method='wilson'). Source: projects/referral-triage-classifier.

A measured recall of 90% on 50 positive cases has a 95% Wilson interval of 78.6% to 95.7%. That’s too wide to tell a good model from a mediocre one. On 500 positives it narrows to 87.1% to 92.3%. The project’s test set had 243 suspected-cancer letters. The model sent 242 of them to a person, and the interval on that recall was still 97.7% to 99.9%.

Write labelling guidelines

Every label needs a definition someone other than the author can apply. Good guidelines:

  • define each label in a sentence, with two or three typical examples;
  • list the known edge cases and the decision for each (“a suspected basal cell carcinoma is routine in this service”);
  • include an “unsure” option, so ambiguous items are flagged instead of guessed;
  • carry a version number, because the guidelines will change after the first round.

For clinical labels, the people applying the guidelines should be people who make the decision in real life, and the guidelines should match the service’s own referral criteria, not a generic list.

Measure agreement between labellers

Have two people label the same sample independently, a few hundred items if you can, and compare. scikit-learn has Cohen’s kappa, which corrects raw agreement for the agreement you’d expect by chance:

from sklearn.metrics import cohen_kappa_score

kappa = cohen_kappa_score(labeller_a, labeller_b)

# For ordered labels, penalise distant disagreements more. Pass the order explicitly:
# without it, string labels are sorted alphabetically.
order = ["Routine", "Urgent", "Suspected cancer"]
weighted = cohen_kappa_score(labeller_a, labeller_b, labels=order, weights="linear")

Then read every disagreement. Most will be guideline gaps, which you fix by updating the guidelines and relabelling. Some will be real ambiguity, which is worth knowing too.

Agreement is also a ceiling. If two clinicians agree on 95% of letters, a model scored against one of them can’t meaningfully reach 99%, and its “errors” will include the other clinician’s calls. The project shows how this looks. I moved 1.5% of the urgency labels on purpose, to stand in for triager disagreement. Of the 12 urgent or suspected-cancer letters the model let through, 9 were those moved labels. In 8 of them a routine letter had been moved up to urgent, so the model was arguably right. With real data there’s no column that tells you which is which. Adjudicating a sample of the model’s confident “errors” is how you find out.

Choose metrics that match the risk

A single accuracy figure hides the thing you need to know. Report:

  • Per-class precision and recall, with the costly class first, and macro averages rather than accuracy when classes are imbalanced.
  • Performance at the operating point. If the output is thresholded, report recall and workload at the threshold you’ll actually use. In the project, choosing the threshold on validation practices gave a fast-track recall of 97.5% with 55.9% of letters skipping the first read.
  • Calibration, if anyone will read a score as a probability or set a threshold on it. An LLM’s verbal confidence (“I’m 90% sure”) is a number it wrote, not a measured probability. Check it against outcomes before using it. See calibration before AUC.
  • Intervals, so a two-point difference on 200 examples isn’t mistaken for progress.
  • The whole-system number: how much human work the feature saves or creates, including the review of its outputs.

For extraction or generation tasks, define what “correct” means per field before scoring (exact match, normalised match, or a rubric applied by a person) and report it per field.

Test for regressions on every change

Prompts get edited, model versions get retired, providers change defaults. Treat each change like a code change: rerun the full test set and compare item by item, not just the headline.

The comparison that matters is between the items that flipped. The project compared two model versions on the same 1,287 test letters. Their accuracy was identical, 95.6%, but each got 10 letters right that the other got wrong. McNemar’s exact test on those 20 letters gave p = 1.0, so there’s no evidence either is better. The headline would have said “no change”. The item-level view says 10 letters now get a different answer, and someone should read them.

statsmodels has the test:

from statsmodels.stats.contingency_tables import mcnemar

# rows: old version right / wrong; columns: new version right / wrong
table = [[both_right, only_old_right], [only_new_right, both_wrong]]
result = mcnemar(table, exact=True)
print(result.pvalue)

A few habits make this work:

  • Pin the model version where the API allows it, and record the version, prompt and settings with every evaluation run.
  • Store every output, not just the scores, so you can diff runs.
  • Measure run-to-run variation. Run the same version twice. If outputs differ between identical runs, that variation is the noise floor for any comparison.
  • Set a gate before you look. For example, “no drop in suspected-cancer recall, and no more than N newly broken items without review.” Decide it before the results come in.

Keep the test set out of the prompts

The quickest way to ruin a test set is to let it leak into what it’s testing.

  • Few-shot examples come from the development set. Never from the test set, and not near-duplicates of test items either.
  • Don’t debug on test items. Pasting a failing test case into the prompt to “fix” it turns the test set into a training set. Move the item to the development set and replace it.
  • Restrict access to the test set and its labels, and keep a log of who ran what against it.
  • Assume public data is contaminated. If your evaluation items are on the internet, a model may have seen them during training. For anything that matters, build the set from your own unpublished data.
  • Refresh it. Inputs drift. Add a fresh sample from recent traffic periodically, and keep the old set for regression comparisons.

Governance for clinical text

In health data, the test set is patient data, and every rule that applies to the feature applies to the evaluation too.

  • Lawful basis and a DPIA. Health data is special category data under UK GDPR. A data protection impact assessment should cover the evaluation as well as the live feature, including where the model provider processes the text and whether inputs are retained or used for training.
  • Minimum necessary. Evaluate on the fields the feature needs. Free-text de-identification is hard to do reliably, so don’t treat it as a guarantee.
  • Synthetic data for development, real data for sign-off. Synthetic letters, like the project’s, are useful for building the pipeline and testing edge cases. They can’t show how the feature performs on real letters. That evidence has to come from governed real data before go-live.
  • Clinical safety. In England, health IT that supports clinical decisions falls under the NHS clinical risk management standards, DCB0129 for manufacturers and DCB0160 for deploying organisations. The evaluation results belong in the safety case, along with the hazards they don’t cover.
  • Medical device questions. Depending on its intended purpose, software that triages patients can qualify as a medical device. The MHRA’s guidance on software and AI as a medical device is where to check.
  • Human oversight by design. Decide which outputs a person must review, and measure how often reviewers overturn the model once it’s live.

A checklist

  1. Write down the decision, the costly error and the fallback.
  2. Sample from real inputs, stratify the rare costly cases, and split into development and test sets.
  3. Write labelling guidelines, double-label a sample, compute kappa, and resolve disagreements.
  4. Pick metrics for the costly class at the real operating point, with intervals.
  5. Pin versions, store outputs, and compare item by item on every change.
  6. Keep test items out of prompts and out of public view.
  7. For clinical text, get governance in place before the first real item is sent anywhere.

If you’re comparing an LLM with a classical model, run both through the same set. Build a text classification baseline before you reach for an LLM covers building the model that sets the bar.