20 May 2026 · 7 min
A generator that memorises its training set reproduces real people. Measurement is not optional.
Synthetic data is often sold as a clean escape from privacy law: generate artificial records that look like the real thing, and because no real person is in them, the rules no longer apply. That story is appealing and, taken at face value, dangerous. Synthetic data can be a genuinely powerful privacy tool — but only when you measure what it actually protects, rather than assuming the word "synthetic" is a magic wand.
The honest picture is more interesting than the marketing. Done carefully, synthetic data lets work proceed that privacy rules would otherwise block. Done carelessly, it leaks the very information it was supposed to hide, while giving everyone a false sense of safety.
Synthetic data is generated by a model that has learned the statistical structure of a real dataset and then produces new records drawn from that learned distribution. Done well, the synthetic set preserves the correlations and patterns that make data useful — so a model trained on it behaves much like one trained on the real thing — without any row corresponding to a real individual.
That last clause is where the subtlety lives. "No row corresponds to a real individual" is a claim about the output, and claims need to be verified, not assumed. The generation process itself determines whether the claim is true.
A generative model can memorise. If it overfits, it may reproduce real records almost verbatim, or produce synthetic points so close to real individuals that those individuals can be re-identified. The dangerous part is that the output still looks synthetic — it passes the eyeball test — while quietly carrying real, sensitive information. A team that generated synthetic data and assumed it was automatically safe would be shipping a privacy breach that no one can see.
This is the failure mode that matters, and it is invisible unless you deliberately measure for it. "It looks fake" is not evidence that it is safe.
The responsible approach treats privacy as a number, not an assumption. We measure how close synthetic records sit to the real individuals they were derived from, and we apply differential privacy during generation where the setting calls for it — a mathematical guarantee that bounds how much any single real person can influence the output. The result is a re-identification risk you can state, defend and put in writing, rather than a hopeful adjective.
Differential privacy is not free: turning up the guarantee tends to blur the data. But it converts privacy from a vibe into a dial you can set deliberately, with the trade-off made explicit rather than stumbled into.
Privacy is only half the equation. Synthetic data that is perfectly private but statistically useless is worthless — you could achieve that with random noise. The other measurement that matters is utility: does a model trained on the synthetic data actually perform when tested against real, held-out data? Only when both numbers are good is the synthetic dataset doing its job.
These two goals pull against each other. More privacy protection tends to cost utility, and more faithful data tends to cost privacy. The engineering is in finding the point on that curve that satisfies both the legal requirement and the practical need, and knowing where that point is requires measuring both, not guessing.
Used honestly, synthetic data is often what makes a stalled project possible. Data that cannot legally cross a border, be shared with a partner, or be used in development can be replaced with a synthetic stand-in that carries the statistical signal without the real records. We have seen initiatives frozen for months on a data-protection question dissolve once the design no longer touches real personal data at all.
The key is that this only works when the privacy claim is measured. The value is real, but it is contingent on doing the measurement the marketing skips.
Synthetic data is not always the right tool. For some tasks the loss of fidelity is unacceptable; for others, simpler techniques like careful anonymisation or access controls solve the problem with less complexity. Reaching for synthetic generation because it sounds sophisticated, when a simpler approach would do, is its own kind of mistake. The tool has to match the problem.
We are candid about this. Sometimes the answer is synthetic data with measured guarantees; sometimes it is a small on-device model so the real data never moves; sometimes it is something plainer. The right choice depends on your data, your constraints and your risk tolerance.
The single practice that separates responsible synthetic data from wishful thinking is documentation: a written statement of the re-identification risk, the utility measured against real data, and the differential-privacy parameters used to get there. That document is what lets your legal team sign off with confidence rather than hope, and it is what turns "it's synthetic, so it's fine" into a defensible position. Without it, you have a claim; with it, you have evidence.
Book a 30-minute call. We will tell you honestly whether we can help.
Book a call