Contents
Figure 1: Two labels, one question — which one actually stops someone re-identifying a person?
Neither one wins by default. Synthetic data and anonymization are different kinds of things, and treating them as interchangeable is the fastest way to ship a dataset you believed was private. One is a legal standard, the other is a technique, and the real answer depends on how each one is built and tested. One caveat first: I'm not a lawyer, so run any real release past your DPO or counsel. This piece sits alongside the differential privacy walkthrough and the privacy-first synthetic data work elsewhere in this series.
Three Terms People Mix Up
Pseudonymization replaces direct identifiers like names or IDs with codes. Pseudonymised data remains personal data, because a key or mapping table can re-link it to a person. It's a useful security measure, not a privacy finish line.
Anonymization is the stricter bar. The GDPR treats a dataset as anonymous when a person can't be identified, directly or indirectly, by any means reasonably likely to be used, considering cost, time, and current and developing technology. Anonymous data falls outside GDPR entirely, which is exactly why everyone wants to claim it.
Synthetic data is newly generated data that mimics the statistics of a real dataset. Whether it counts as anonymous isn't automatic — anonymization is a legal standard rather than a statistical one, so labeling data "synthetic" doesn't settle the question.
How Traditional Anonymization Fails
Classic anonymization modifies real records by deleting, generalizing, or replacing identifying details. The workhorse is k-anonymity: group records so each combination of identifying attributes appears for at least k people. Its cousins, l-diversity and t-closeness, add requirements on variation and distribution similarity within those groups.
The weaknesses are well documented. K-anonymity remains the de facto minimum standard according to the European Medicines Agency, yet it reduces data utility, is vulnerable to several forms of attack, and is computationally hard to perform. A deeper problem is deciding which attributes count as quasi-identifiers, since you have to anticipate what an attacker might link against, today and later.
That linkage risk takes three forms: individualization (isolating one person), correlation (linking datasets about the same individual), and inference (deducing information about someone). Because modern technology makes cross-source linking easier every year, guaranteeing zero re-identification risk is close to impossible. Notice what this means: a k-anonymized table still contains real people's generalized records, so any auxiliary dataset becomes a potential key.
How Synthetic Data Changes the Picture
Properly generated synthetic data has a structural advantage here. There's no direct association or one-to-one mapping between synthetic records and an individual's data, so an attacker has no row to link against. That removes the most common attack path against anonymized tables.
But the risk doesn't vanish, it moves. Even without exact identifiers, machine learning models can encode patterns from the original data, creating subtle correlations that allow re-identification in poorly generated datasets. Regulators have noticed: with generative AI, the risk no longer comes only from the dataset itself but also from the behaviour of the model that generates it, a concern identified by ENISA and NIST. This is the same membership inference and memorization problem covered earlier in this series.
There's also an unavoidable tradeoff. Higher utility generally means lower anonymity, because a dataset that closely replicates the original makes identifying individuals easier. A synthetic dataset faithful enough to be useful is, almost by definition, leaking some real signal.
What Regulators Actually Say
The legal picture is still settling. Creating synthetic data from real personal data counts as processing under GDPR, and whether synthetic data itself remains personal data is a complex, unresolved question. The EU AI Act points toward treating it as non-personal, since Article 59 lets regulatory sandboxes use synthetic or anonymised data instead of personal data where it can meet the objective, but the Act doesn't set privacy standards that synthetic data must meet.
The EDPB's position is cautious. Data counts as anonymous only if re-identification stays reasonably impossible considering current and future technical means, which calls for a case-by-case assessment, especially with sensitive, extensive, or highly granular source data. In practice, an organisation can't assume a dataset falls outside GDPR just because it's labeled synthetic, and should document re-identification risks, test the generating models' robustness, control the origin of training data, and assess vendor safeguards.
So Which Protects Privacy Better?
Rough ranking, assuming competent execution of each:
- Differentially private synthetic data sits at the top. DP gives a mathematical bound on what any output reveals about one person, and that bound holds even against attackers with auxiliary information. K-anonymity offers no equivalent guarantee, which is the core reason DP was invented. The previous piece in this series covers the mechanics and costs.
- Non-DP synthetic data from a well-validated generator comes next. It removes direct record linkage but offers no formal bound, so its safety rests entirely on empirical testing.
- K-anonymity and its variants follow, since they keep real (generalized) records exposed to linkage attacks and degrade utility as protection tightens.
- Pseudonymization is last, because it's reversible by design and still personal data.
That ranking comes with caveats. Poorly built DP, with budget leaked through data-dependent preprocessing, can protect less than its ε suggests. A sloppy synthetic generator that memorizes training rows can be worse than careful aggregation. And for small, simple releases, like a handful of published statistics, straightforward aggregation with suppression of small cells may be simpler and more transparent than any generator.
How to Test Whatever You Choose
Don't trust the label, test the output. A reasonable validation set includes:
- Distance to closest record: for each synthetic row, find its nearest real row. Rows that are near-copies signal memorization.
- Membership inference testing: try to distinguish training records from held-out records using the synthetic output. Success rates well above chance indicate leakage.
- Attribute inference testing: check whether an attacker with partial knowledge about a person can fill in sensitive fields more accurately using your data.
- Singling-out checks: look for rare, unique combinations that identify one person, in the synthetic data or the anonymized table alike.
The first check takes about ten lines with scikit-learn already installed, and it catches the most common failure — a generator that copied rows:
import pandas as pd
import numpy as np
from sklearn.neighbors import NearestNeighbors
real = pd.read_csv("real_data.csv")
synth = pd.read_csv("synthetic_data.csv")
nn = NearestNeighbors(n_neighbors=1).fit(real)
distances, _ = nn.kneighbors(synth)
print(np.percentile(distances[:, 0], [1, 5, 50]))
Run it as python audit.py and look at the low percentiles: a cluster of near-zero distances means the synthetic data contains near-copies of real records, no matter what the label on the dataset says.
Document the results. That documentation is what a DPO or regulator will actually want to see.
A Practical Decision Guide
- Internal analytics where you need to link back to people: use pseudonymization plus strong access controls, and treat the data as personal data throughout.
- Sharing data with partners or vendors for model development: use DP synthetic data where the use case tolerates the utility loss, plus empirical leakage testing.
- Public release: prefer DP mechanisms or heavily aggregated outputs, since public release means assuming a motivated attacker with outside data.
- Rare, high-sensitivity records, such as unusual diagnoses or outlier transactions: be most cautious here, since uniqueness is what both anonymization and synthesis struggle to hide. Consider keeping analysis inside a controlled environment instead of releasing data at all.
Common Mistakes
- Calling pseudonymized data anonymous, which misclassifies personal data and creates real compliance exposure.
- Assuming "synthetic" means "anonymous" without testing the generator or documenting residual risk.
- Evaluating only the synthetic dataset and ignoring the generator, even though regulators now treat the model's behavior as part of the risk surface.
- Optimizing utility without checking the tradeoff, then discovering later that the high-fidelity dataset sits close to the original records.
- Treating the privacy assessment as one-time, when new auxiliary datasets and attack techniques keep changing what "reasonably likely" means — and when governance expectations move with it.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Privacy Law Fundamentals | the vocabulary this whole debate runs on: what counts as personal data, when pseudonymization is still processing, and where the anonymization line is drawn. | View on Amazon |
![]() |
Anonymizing Health Data | practical case studies of k-anonymity variants, quasi-identifier selection, and the linkage attacks that break them. | View on Amazon |
![]() |
Statistical Disclosure Control | the methods behind safe publication: masking, aggregation, small-cell suppression, and the tradeoffs a release owner has to document. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
Is synthetic data anonymous under GDPR?
Not automatically. Anonymization is a legal standard, not a statistical one, so labeling data "synthetic" doesn't settle the question. Creating synthetic data from real personal data counts as processing under GDPR, and whether the output itself remains personal data is a complex, unresolved question. The EDPB's position is that data counts as anonymous only if re-identification stays reasonably impossible considering current and future technical means, which calls for a case-by-case assessment rather than a label.
Is pseudonymization the same as anonymization?
No. Pseudonymization replaces direct identifiers like names or IDs with codes, but pseudonymised data remains personal data because a key or mapping table can re-link it to a person. It is a useful security measure, not a privacy finish line. Anonymization is the stricter bar where a person can't be identified, directly or indirectly, by any means reasonably likely to be used, and anonymous data falls outside GDPR entirely.
Which protects privacy better, synthetic data or k-anonymity?
Assuming competent execution: differentially private synthetic data sits at the top because it gives a mathematical bound that holds even against attackers with auxiliary information; non-DP synthetic data from a well-validated generator comes next, removing direct record linkage but offering no formal bound; k-anonymity follows, since it keeps real generalized records exposed to linkage attacks while degrading utility; pseudonymization is last because it's reversible by design. Poorly built versions of any of these can reorder the ranking.
How do I test whether my synthetic data is private?
Don't trust the label, test the output. Check distance to the closest real record for signs of memorization, run membership inference tests to see whether training records are distinguishable from held-out ones, try attribute inference with partial knowledge about a person, and look for singling-out combinations that identify one person. Document the results, since that documentation is what a DPO or regulator will actually want to see.
Wrapping Up
Anonymization is a legal outcome, synthetic data is a technique, and pseudonymization is a security measure, so asking which "protects better" really means asking how each is executed and verified. Differentially private synthetic data gives the strongest formal footing, careful synthetic data without DP removes record-level linkage but relies on testing, and traditional anonymization stays exposed to linkage while costing utility.
Can any of them promise zero risk? No, and the better frameworks say so openly. What you can do is pick the method whose guarantees match your threat model, test the output instead of trusting the label, and document the reasoning before someone asks for it.


