Sam Austin on October 8, 2026

Differential Privacy in Synthetic Data: Mathematical Privacy Guarantees

Differential Privacy in Synthetic Data: Mathematical Privacy Guarantees
Contents

Code on a screen representing the mathematical guarantees behind differential privacy

Figure 1: Synthetic is not private — the guarantee has to live in the mechanism, not the label on the file

Calling a dataset "synthetic" doesn't make it private. A generator that memorizes its training data can leak individual records just as surely as the original table, which is why membership inference attacks keep showing up in this series. Differential privacy (DP) is the one approach that gives you a mathematical guarantee instead of a hopeful assumption. It also comes with real costs and a few subtle traps, so let's go through both, alongside the privacy-first synthetic data work covered elsewhere in this series.

What the Guarantee Actually Says

Differential privacy makes a promise about a mechanism — the algorithm that produces your synthetic data — not about any particular output. Picture two worlds: one where your dataset includes a specific person's record, and one where it doesn't. A DP mechanism guarantees an observer can barely tell which world produced the output, so no individual's presence or absence meaningfully changes what you release.

Two parameters quantify "barely." Epsilon (ε) is the privacy loss budget: smaller means the two worlds look more alike and the guarantee is stronger. Delta (δ) is a small failure probability, the chance that the ε bound doesn't hold, which you want far smaller than one over your dataset size. Formally, a mechanism satisfies (ε, δ)-DP when, for any two datasets differing in one person's record, the probability of any output changes by at most a factor of e^ε, plus δ.

That framing exposes the fundamental tradeoff. Any release that meaningfully preserves the original data's statistics leaks some information, so privacy and utility pull against each other, and ε is the dial between them.

Why DP Gives You More Than "It's Synthetic"

DP has a property that makes it especially natural for synthetic data: post-processing is free. Once a mechanism produces differentially private statistics or model parameters, anything you derive from them without touching the original data again remains DP. You can sample a million synthetic rows, or ten million, without spending additional privacy budget. The guarantee attaches to the noisy intermediate object, not to how many records you pull out of it.

The catch is that every decision that peeks at the real data costs budget, including choices people often treat as free, like which columns to model or how to bin a numeric field.

The Two Families of DP Synthetic Data

Research on DP synthetic data splits into two broad families, and the NIST Differential Privacy Synthetic Data Challenge pitted them against each other.

Marginal-based methods work in three steps: select a set of low-dimensional marginals (counts over one or two columns), measure them with calibrated noise, then fit a model and sample synthetic records that match those noisy measurements. MST, the winner of the 2018 NIST challenge, picks pairwise marginals using a maximum spanning tree over mutual information scores, measures them with Gaussian noise, then fits a graphical model through Private-PGM. AIM extends MST by choosing marginals adaptively, measuring whichever one would reduce model error most and spending its privacy budget on the most informative statistics first, which gives it higher utility than MST at the same ε. PrivBayes and PrivSyn follow related ideas using Bayesian networks and model-free marginal synthesis.

Deep generative methods apply DP to neural training itself, typically through DP-SGD, which clips each example's gradient and adds noise during training. DP-GAN variants, DP versions of TVAE, and newer approaches like fine-tuning language models as DP tabular generators all live here. They handle complex, high-dimensional data more flexibly, at the price of heavier compute and a harder time proving tight guarantees.

For most tabular problems, the marginal-based family earns its reputation. MST and AIM consistently show strong privacy-utility tradeoffs, appear in popular libraries, and even supported a UK census data release. One caveat: marginal-based methods work on discretized data, so they tend to struggle on datasets with many numerical columns, where binning loses detail.

A Minimal Example

SmartNoise's snsynth package ships MST and several other DP synthesizers behind one API. Install with pip install smartnoise-synth. The exact arguments change between releases, so check the current docs before running this:

import pandas as pd
from snsynth import Synthesizer

df = pd.read_csv("sensitive_data.csv")

synth = Synthesizer.create("mst", epsilon=1.0)
synth.fit(df, preprocessor_eps=0.5)

synthetic_df = synth.sample(10_000)

Notice the preprocessor_eps argument. When a library infers things like column bounds or category lists from the raw data, that inference consumes privacy budget, which is exactly the hidden cost mentioned above. If you can define bounds and categories from public knowledge instead, you spend your whole budget on the actual measurements.

Choosing Epsilon Honestly

Nobody can hand you the "right" ε, because it depends on your threat model, your data's sensitivity, and how much utility you can sacrifice. A few honest points help.

Smaller ε means stronger guarantees and noisier data. Benchmarks regularly report results at larger values like ε=10 or 20 for comparability with earlier work, and researchers explicitly note that such values correspond to weaker theoretical guarantees, though DP mechanisms can still reduce practical attack success compared with non-DP synthetic data. That's useful context, but it doesn't make ε=10 a strong guarantee. Treat it as "better than nothing," not "private."

Also, an ε value alone doesn't describe a guarantee. NIST's SP 800-226 guidance on evaluating DP guarantees emphasizes that a guarantee is defined by both ε and its other parameters, so you should report all original privacy parameters, including δ and what counts as one individual's data, rather than quoting ε by itself. Two systems advertising the same ε can protect people very differently.

Does the Theory Hold Up in Practice?

Encouragingly, yes, at least for the best-studied methods. Researchers recently built an auditing framework that measures empirical privacy leakage against theoretical bounds. For MST and AIM at (ε,δ)=(1, 10⁻²), the measured leakage came out at roughly μ≈0.43 against an implied theoretical value of 0.45, a small gap between theory and practice. In plain terms, the guarantee isn't just paper math: the empirical attack results track the theoretical bound closely. The same work points out that earlier audits reporting a single (ε,δ) setting could mislead, which is another reason to examine the full privacy tradeoff rather than one number.

The Costs You Should Expect

DP isn't free, and the costs show up in several places. Utility drops as ε shrinks, since more noise means blurrier statistics, with the biggest losses on rare categories and fine-grained correlations. That matters a great deal for the rare-event detection problems covered elsewhere in this series, where the minority class is exactly the signal DP noise tends to wash out.

Fairness suffers too. Recent benchmarking found that DP alone generally reduces utility and can exacerbate group disparities, though fairness interventions can partially recover fairness on DP synthetic data. Small subgroups carry less signal relative to the noise, so they degrade first. If your downstream model affects people differently by group, test subgroup performance on the DP synthetic data explicitly.

Common Pitfalls

  • Treating the output as the thing that's private rather than the mechanism. A DP guarantee covers the whole algorithm, including every data-dependent choice you make along the way.
  • Letting data-dependent preprocessing or hyperparameter tuning sneak past the budget. Selecting columns, bins, or model settings by looking at the real data spends privacy you haven't accounted for, so tune on public or separate data where possible.
  • Releasing many synthetic datasets from repeated runs without composing the budget. Each fresh mechanism run on the same data consumes additional ε, and the losses add up.
  • Quoting ε without δ or the unit of privacy, which makes the guarantee impossible to compare or evaluate.
  • Reading ε=10 or 20 as strong protection when it's a weak theoretical guarantee that happens to reduce attack success empirically.
  • Assuming DP replaces governance. It reduces disclosure risk, but legal and compliance review still matters, especially because regulators are still settling what evidence validates a synthetic dataset.
CoverBookDescriptionGet it
Cover of “The Algorithmic Foundations of Differential Privacy” The Algorithmic Foundations of Differential Privacyby Cynthia Dwork and Aaron Roth the canonical treatment: the definition, composition, and the proofs behind the ε and δ accounting this whole piece depends on. View on Amazon
Cover of “Differential Privacy: From Theory to Practice” Differential Privacy: From Theory to Practiceby Ninghui Li et al. the bridge from definitions to working systems, including how privacy budgets get spent by the practical choices inside a pipeline. View on Amazon
Cover of “Privacy Engineering: A Developer's Guide” Privacy Engineering: A Developer's Guide the governance layer that a formal guarantee still doesn't cover: threat models, deployment decisions, and what regulators expect to see documented. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What does differential privacy actually guarantee?

It guarantees a bound on how much the presence or absence of any single individual's record can change the probability of any output from the mechanism. Formally, for any two datasets differing in one person's record, the probability of any output changes by at most a factor of e^ε, plus a small failure probability δ. The promise attaches to the whole algorithm, including every data-dependent choice inside it, not to one particular synthetic file.

What epsilon value should I use?

There is no universal right answer, because ε depends on your threat model, your data's sensitivity, and how much utility you can sacrifice. Smaller ε means stronger protection and noisier data. Values around ε=1 to 2 are commonly treated as a meaningful guarantee for published analyses, while ε=10 or 20 reported in benchmarks is a weak theoretical guarantee that happens to reduce attack success empirically — better than nothing, not "private".

Does differential privacy make synthetic data safe to release publicly?

It makes disclosure risk quantifiable rather than hopeful, but only if the entire pipeline satisfies DP, including preprocessing, column selection, and hyperparameter tuning. DP reduces the risk that any individual's record leaked into the output, yet it does not replace governance: legal and compliance review still matters, especially because regulators are still settling what evidence validates a synthetic dataset.

Do I spend more privacy budget by generating many synthetic rows or repeated datasets?

Generating more rows from one DP output is free, because post-processing of a differentially private result stays differentially private no matter how many records you sample. Running the mechanism again on the same raw data is not free: each fresh run consumes additional budget, and composition rules make the losses add up across releases, so count every run against your total ε.

Wrapping Up

Differential privacy turns "we think this synthetic data is safe" into a quantified, auditable guarantee on the generating mechanism, with ε as the dial between privacy and utility. Marginal-based methods like MST and AIM give the best tradeoffs for most tabular data, deep DP generators cover harder structure at higher cost, and recent audits suggest the theoretical guarantees hold up empirically. The costs are real, though: lost utility, worse performance on small groups, and a budget that leaks away through any data-dependent choice you forget to count.

Will DP make your synthetic data perfectly safe and perfectly useful? No, the tradeoff is fundamental, and no tuning removes it. But it's the only approach that lets you state precisely how much you're giving up on each side, which beats crossing your fingers and calling the data synthetic.

What are You Looking For?

esc