Behnam Analytics

Writing Data analysis & statistics

Synthetic health data: what it's for and where it stops

What synthetic data is good for (teaching, pipeline testing, demos, sharing structure), what it can't support, how the main approaches differ, and why synthetic isn't automatically anonymous.

Behnam Ebrahimi 8 min read

Synthetic data is data that was generated rather than recorded. In health analytics it gets used for two very different jobs. One is standing in for real data so people can build, test and teach without access to patient records, which it does well. The other is producing findings about real patients, which it can’t do, however realistic it looks.

Most of the trouble comes from mixing the two up. This article sets out where the line falls, compares three ways of synthesising a table (with numbers), explains why “synthetic” doesn’t mean “anonymous”, and describes how the projects on this site generate their data and what is planted in it.

What it’s good for

Teaching. Someone learning SQL, DAX or survival analysis needs data with the right shape and the right problems: skewed waits, missing codes, duplicate referrals. They don’t need real patients, and a training environment shouldn’t hold them.

Testing pipelines. Most pipeline bugs are about structure, not values: a column that changes type, a code that arrives in lower case, a date in the future, a join that fans out. Synthetic data can contain every one of those on purpose, at whatever volume you need, and a test suite can run against it anywhere.

Demos and portfolios. A dashboard or a model can be shown in full, including the data behind it, without an information governance conversation. Every project on this site works this way.

Sharing structure before access. Research teams often wait months for access to real data in a secure environment. A synthetic extract with the same tables, columns and codes lets them write and debug their code first, so the time with real data goes on analysis. NHS England’s artificial data pilot does exactly this for Hospital Episode Statistics: sample and full-size artificial versions of the A&E, admitted patient care and outpatient datasets, for developing and testing code.

Testing methods against a known truth. When you plant an effect, you know the right answer. That makes simulated data the best way to check whether a method finds what it should. The forecasting, SPC and funnel plot examples on this site all rely on it.

What it can’t do

Support clinical or operational conclusions. Whatever a synthetic table says about real patients, it says because the generator put it there. If the generator was fitted to real data, you are looking at a smoothed, partial echo of that data. If it wasn’t, you are looking at your own assumptions.

Validate a model. A model’s performance on synthetic data tells you about the generator, not about patients. The numbers below show the same model scoring anywhere from 0.595 to 0.685 depending only on how the training data was synthesised.

Answer anything that depends on relationships the generator didn’t keep. Many generators reproduce each column’s distribution well and the relationships between columns badly or not at all. Subgroups, interactions and rare combinations are the first things to go, and they are often the reason for doing the analysis.

The NHS England artificial data is a clear example of drawing this line on purpose. Its documentation says each column is aggregated independently and fields are generated at random, so no relationships between fields are preserved and a young patient can have a geriatric diagnosis. It is explicit that the artificial data is not for analysis. That is a design choice, not a flaw: the lack of relationships is what protects confidentiality, and it is still ideal for writing code.

Three ways to synthesise a table

To show the trade-offs with numbers, I simulated a small admissions table (age, sex, deprivation quintile, number of long-term conditions, length of stay and a 30-day readmission flag) to stand in for a real extract. It is itself synthetic, so no real patients are involved anywhere. I planted clear relationships: older patients have more conditions, longer stays and more readmissions.

I then synthesised 3,000 rows of it three ways, and held back 2,000 source rows the generators never saw:

  1. Independent columns. Sample each column separately from its own distribution. This is the approach the NHS England artificial data takes.
  2. Gaussian copula. Keep each column’s own distribution, and join them using the rank correlations between columns. This is a common statistical approach, and you can write it in about ten lines.
  3. Noisy copies. Copy real rows and nudge age and length of stay by up to one. It is tempting because the result looks so realistic.
Rank correlations in the source data and in each synthetic copySpearman correlation; independent columns lose every relationship
Data table
Data setAge and LTCsAge and length of stay
Source data0.440.51
Independent columns0.00-0.03
Gaussian copula0.430.48
Noisy copies0.440.48

Synthetic data. Source: projects/article-examples/stats/synthetic_health_data.py.

Source Independent Copula Noisy copies
Mean age 65.6 65.3 66.7 65.4
Readmission rate 18.5% 19.2% 19.3% 19.1%
Rank correlation, LTCs and readmission 0.17 0.00 0.08 0.15
AUROC of a model trained on it, tested on held-out rows 0.684 0.595 0.675 0.685

Every method got the single-column summaries about right. That is the trap: a synthetic table can pass every check on means and distributions and still be useless for analysis.

Independent columns destroyed every relationship, as designed, and a readmission model trained on them lost most of its ability to rank patients (an AUROC of 0.595, against 0.684 for a model trained on the source). The copula kept the correlations between continuous columns (0.43 for age and conditions, against 0.44 in the source) but halved the link between conditions and readmission. Relationships involving a yes/no outcome are exactly what a Gaussian copula handles worst. The noisy copies kept everything, which should worry you more than it reassures you.

Synthetic isn’t automatically anonymous

The ICO’s guidance on synthetic data is direct about this: whether synthetic data is anonymous depends on whether the personal information it was modelled on can be inferred from it. It notes that some generation methods have been shown to be vulnerable to membership inference, model inversion and attribute disclosure, and that outliers, the records with unusual combinations, carry the most risk.

The mechanism is simple. The more closely synthetic data mimics the real data, the more useful it is and the more likely it is to reproduce real records. Flexible generative models (GANs, variational autoencoders, large language models) can memorise rare records from their training data, and a memorised record looks like any other synthetic row.

A first disclosure check is to count synthetic rows that match a source row, either exactly or nearly. The count only means something next to a baseline. Some matches happen by chance, especially with few, coarse columns, so compare against held-out source rows that no generator saw:

Rows that match a record in the 3,000-row sourceHeld-out source rows, never shown to a generator, give the chance level
Data table
Data setExact matchNear match (age and stay within 1)
Held-out rows (chance)9%48%
Independent columns7%38%
Gaussian copula10%46%
Noisy copies21%100%

Synthetic data. Source: projects/article-examples/stats/synthetic_health_data.py.

With only six coarse columns, 8.6% of held-out rows exactly match some source row, and 48.2% nearly match one, purely by chance. Independent columns and the copula sit at or near that level: 6.8% and 9.9% exact. The noisy copies are at 20.6% exact, and every one of them is within a year of age and a day of stay of a source record. That is the data itself with a thin disguise. And 88.7% of source rows were unique on these six columns, so a near copy of a unique row points straight at one person.

Real extracts have more columns and dates, which makes chance matches rarer and genuine copies easier to spot. It also makes more records unique, so each copy discloses more.

If synthetic data derived from patient records is going to leave a secure environment, treat it as a disclosure control decision: check matches against a holdout baseline, look hard at outliers and rare combinations, and involve whoever signs off your data releases. Differential privacy gives formal guarantees, at a cost in utility that the ICO guidance also notes.

How the projects on this site make theirs

Nothing on this site is fitted to real patient data. Every dataset is simulated from rules: I write down how the process works (arrivals, weekly and seasonal cycles, how risk depends on age and conditions), generate records from those rules with a fixed random seed, and plant specific patterns for the analysis to find. Because no real records go in, there is nothing to disclose. The price is that realism depends entirely on the rules. The data can surprise you only in the ways you allowed it to.

The rule I follow is to document what is planted, so readers can check each method against the truth:

  • ED demand forecasting: daily attendances built from trend, day of week, a winter cycle with a random respiratory wave, bank holidays, heatwaves and data-feed outages, with persistent noise on top. The download keeps the planted expected value for every day next to the noisy count, so you can see how much of a forecast miss was the model and how much was noise.
  • Readmission risk model: patient histories with non-linear effects and interactions the model doesn’t include, missing values, a trend over time, and a deliberate leakage trap.
  • RTT waiting list semantic model: pathways with long-tailed waits, a capacity squeeze through one winter and validation sprints.
  • Outpatient analytics pipeline: an extract with planted data-quality problems, such as a site using legacy outcome codes for part of the period, so the data tests have something real to catch.
  • Bed occupancy simulation: decisions to admit with hourly and weekly cycles, skewed stays in three patient groups, and fewer discharges at weekends, each a named constant so every result traces back to its inputs.
  • Referral triage text classifier: referral letters built from templates with abbreviations, typos, negation and a share of deliberately moved labels, so every error has a known cause.
  • The articles: a step change and a disrupted week in the SPC example, three planted outlier practices in the funnel plot example, and a U-shaped effect and an interaction in the model comparison.

Each project documents its planted patterns in full alongside its code.

A checklist

When someone hands you synthetic data, or you make some:

  1. Ask how it was made and whether real records went in. Rule-based, fitted to real data, or copied with noise are very different things.
  2. Check which relationships survive before you analyse anything. Compare correlations and a simple model against the source, not just column summaries.
  3. Never report findings from it as if they describe patients. Label it as synthetic on every chart and file.
  4. If it came from real records, run a disclosure check: exact and near matches against a holdout baseline, and a look at the rarest rows.
  5. Write down what is planted, if you simulated it. A dataset with a known truth is only useful if people know what the truth is.

Reproduce

The comparison comes from projects/article-examples/stats/synthetic_health_data.py, which simulates the source table, runs the three generators, and prints the fidelity and disclosure checks:

uv run python projects/article-examples/stats/synthetic_health_data.py