Sam Austin on October 11, 2026

Bias in Synthetic Data: How Generation Methods Can Amplify Unfairness

Bias in Synthetic Data: How Generation Methods Can Amplify Unfairness
Contents

Abstract visualization of skewed data patterns representing bias in generated datasets

Figure 1: Rebalanced counts are not the same thing as fair outcomes

Synthetic data gets pitched as the cure for biased datasets. Generate more records for the under-represented group, rebalance, retrain. The evidence says that pitch is half right, and the other half is where people get hurt.

Here's a concrete example from a diffusion-based tabular generator study. Adding Tab-DDPM synthetic data improved fairness in binary classification overall, yet for one classifier (Gaussian Naive Bayes) the statistical parity difference rose as synthetic samples were added, which the authors take as a sign that synthetic data can make bias worse. Same generator, same dataset, opposite outcome depending on the downstream model — the same recurring finding as the text and image augmentation articles: whether synthetic data helps depends on the task, not just the generator.

This article covers how bias enters synthetic data, how generation methods can amplify it, and the checks that tell you which way your dataset is going. It also brings fairness into the fidelity–utility–privacy framework used across this series' evaluation discussions. A dataset can pass all three and still be unfair. The sentence worth keeping: a model trained on debiased synthetic data may appear fair while operating in a biased world.

Three Ways Bias Gets In

1. The generator learns the real data's bias. A generator that faithfully reproduces its training distribution reproduces the biases in it. High fidelity, the goal of the SDV and CTGAN articles, is the problem here, not the fix.

2. The generator brings its own bias. This is specific to LLM-based generation: LLMs reflect biases present in their training data, so when they generate synthetic data for training, they can propagate and amplify those biases — a phenomenon the literature calls bias inheritance. The mechanism is concrete: a model fine-tuned on a mix of real and synthetic data adopts, and often amplifies, the social biases contained in the synthetic portion. The defining study measures it with a bias ratio — the share of augmentation data in the full training set — using synthetic data generated by LLaMA 3.1 and GPT-4o-mini.

3. Feedback loops across generations. The fundamentals article covers model collapse; fairness has its own version. Models trained on data from previous-generation generators have been shown to amplify bias over time, and recursive training on synthetic data without periodic re-grounding in real data may drift. In multi-round classification experiments, performance declined across all demographic groups over successive rounds.

LLM Generators Deserve Extra Suspicion

For tabular data, LLMs often generate rows from a handful of real examples placed in the prompt. One GPT-4o study set out to measure how far LLM-generated tabular data is shaped by social biases and stereotypes, and how the characteristics of the prompt data — the number and makeup of in-context examples — influence the result. The practical lesson: the examples you put in the prompt are part of your fairness pipeline. If your five demonstration rows skew toward one group, your thousand generated rows probably will too.

The chatbot training article raises the same concern for dialogue: if the model's default idea of a "nurse" or an "engineer" carries a gender, your persona set will inherit it. That's an inference from the mechanism rather than a cited finding, so audit persona distributions instead of assuming them.

Class Imbalance Is Not Group Imbalance

The SMOTE article fixes imbalance in the target class. Fairness problems often come from imbalance in a protected attribute, and the two frequently overlap: class imbalance and group imbalance commonly coincide in real tabular datasets, both can undermine utility and fairness, and relatively few methods address them together, most relying on interpolation-style oversampling. A comparative study across four datasets found that modern generative models are effective for mitigating this kind of bias.

SMOTE can balance classes and still leave a group badly under-represented inside the minority class. If you only check class balance, you'll miss it.

What Fairness-Aware Generation Looks Like

A growing family of generators tries to build fairness into generation itself — most studies in this area have focused on GAN- or diffusion-based frameworks. Named examples include TabFairGAN for fair tabular GANs, DECAF for causally-aware generative networks, and FairCauseSyn for causally fair LLM-augmented generation. FairDiffuseVQVAE does it at sampling time inside a tabular diffusion model.

Simpler options exist too:

  • Counterfactual edits. Swap demographic, sentiment, or context attributes in real records to create matched pairs; this works best for text-heavy tabular and conversational data. Matched pairs also give you a direct test: if changing only the group attribute changes the model's output, you've found the bias.
  • Targeted generation for under-represented slices. LLM expansion, GAN or diffusion sampling, and rule-based templates can produce records for those slices to rebalance data before fine-tuning.
  • Reweighting. The Tab-DDPM study evaluated sample reweighting as a bias mitigation approach, alongside the synthetic augmentation.

The Limits: Fairness Theater

The FairDiffuseVQVAE authors state the risks bluntly, and they apply to every method above:

  • A model trained on debiased synthetic data may appear fair while operating in a biased world, so the generator can mask bias in a downstream pipeline.
  • Fairness metrics computed on a classifier trained on synthetic data describe that classifier on the original test distribution, not the fairness of real decisions about real people.
  • Elevated distance-to-closest-record scores can create a false sense of privacy. DCR measures one thing, and passing it doesn't mean the rest is fine.
  • Synthetic data shouldn't be treated as a stand-alone fairness solution; pair it with qualitative audits and classifier-side fairness interventions.

Fairness criteria can also conflict with each other, so "fair" has to be defined for your application before you measure anything.

Measuring It: A Fourth Pillar

Add fairness metrics next to fidelity, utility, and privacy. Two common ones: statistical parity difference (the gap in positive-prediction rates between groups) and disparate impact (the ratio of those rates — unintentional bias where predictions produce different error rates or outcomes across groups defined by protected attributes such as race, sex, religion, or age).

import numpy as np

def fairness_report(y_pred, group, privileged):
    y_pred, group = np.asarray(y_pred), np.asarray(group)
    p_priv = y_pred[group == privileged].mean()
    p_unpriv = y_pred[group != privileged].mean()
    return {
        "statistical_parity_difference": p_unpriv - p_priv,
        "disparate_impact": p_unpriv / p_priv if p_priv > 0 else float("nan"),
    }

Then run the dose-response test the Tab-DDPM study used: add generated data in increments to the original dataset and compare balanced accuracy and fairness metrics against the unaugmented baseline.

for n_syn in [0, 500, 1000, 2000, 4000]:
    X_aug = np.vstack([X_real, X_synth[:n_syn]])
    y_aug = np.concatenate([y_real, y_synth[:n_syn]])
    model.fit(X_aug, y_aug)
    print(n_syn, fairness_report(model.predict(X_test), g_test, privileged=1))

Do this with several classifier types, not one. The Naive Bayes result above shows why: a generator can look fair under one model and unfair under another. Always evaluate on a real, held-out test set — the TSTR discipline from the NLP synthetic data article.

Wiring It Into the Rest of the Pipeline

  • Great Expectations. Add expectations on protected-attribute shares and on outcome rates per group, so a generator run that quietly shifts them fails validation.
  • Model monitoring and slice-based evaluation. Track fairness metrics per group in production, not just overall accuracy.
  • Re-grounding. Keep real data in the training mix at every round, as the feedback-loop finding implies.

Common Mistakes People Make

  • Assuming more synthetic data for the minority group automatically improves fairness. The Naive Bayes result shows it can make things worse, and the direction depends on the model.
  • Checking class balance and ignoring group balance. The two overlap but aren't the same, and SMOTE-style fixes address only one.
  • Ignoring the prompt when using LLM generators. Biased in-context examples propagate into the output.
  • Evaluating fairness on synthetic data only. Test on real data. Metrics from a synthetic test set describe the generator, not the world.
  • Treating one fairness metric as the answer. Pick metrics that match your application, and don't assume they all agree.
  • Recursing without real data. Each round of synthetic-on-synthetic training risks compounding drift.
CoverBookDescriptionGet it
Cover of “Fairness and Machine Learning” Fairness and Machine Learningby Barocas, Hardt & Narayanan the definitions behind statistical parity, disparate impact, and why fairness criteria conflict. View on Amazon
Datasheets for Datasets documentation discipline for any dataset, synthetic or real: composition, collection process, and recommended uses. View on Amazon
Cover of “Human-in-the-Loop Machine Learning” Human-in-the-Loop Machine Learning the qualitative audits and review loops that pair with synthetic generation when fairness matters. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

How does bias get into synthetic data?

Three routes. The generator learns the real data's bias — a faithful generator faithfully reproduces whatever unfairness is in its training distribution, so high fidelity is the problem, not the fix. The generator brings its own bias — LLMs reflect biases in their training data and propagate them into generated records, a phenomenon called bias inheritance. And feedback loops amplify bias across generations when models are trained recursively on previous generations' synthetic data without re-grounding in real data.

Can adding synthetic data for a minority group make fairness worse?

Yes. In a diffusion-based tabular generator study, adding Tab-DDPM synthetic data improved fairness in binary classification overall, yet for Gaussian Naive Bayes the statistical parity difference rose as synthetic samples were added. Same generator, same dataset, opposite outcome depending on the downstream model — which is why fairness has to be measured with several classifier types, not one.

What fairness metrics should I check on synthetic datasets?

Two common ones: statistical parity difference (the gap in positive-prediction rates between groups) and disparate impact (the ratio of those rates — unintentional bias where predictions produce different error rates across protected groups such as race, sex, religion, or age). Compute them on real held-out data with several classifier types, and run a dose-response test: add generated data in increments and compare against the unaugmented baseline.

Does passing fidelity, utility, and privacy checks make a synthetic dataset fair?

No — that's the gap this article covers. A model trained on debiased synthetic data may appear fair while operating in a biased world, and fairness metrics computed on a classifier trained on synthetic data describe that classifier on the original test distribution, not the fairness of real decisions about real people. Fairness needs its own pillar, plus qualitative audits and classifier-side interventions.

Wrapping This Up

Synthetic data can reduce bias, by filling gaps and enabling counterfactual tests, and it can amplify bias, by inheriting it from real data, from the generator, and from feedback loops. Which one happens depends on the generator, the prompt, the downstream model, and how you measure. The fixes are procedural: audit group balance as well as class balance, keep real data in every round, test fairness across multiple models on real held-out data, and treat synthetic data as one tool within a broader fairness process.

Remember that passing fidelity, utility, and privacy checks says nothing about fairness, and a fair-looking result on synthetic data can mask a biased real-world pipeline. This closes out the synthetic data arc's evaluation story: the earlier articles gave the fidelity–utility–privacy checks, and this adds the fourth pillar that decides whether the other three matter for real people.

Now take any synthetic dataset you've generated in this series and compute the share of each protected group, plus the outcome rate per group, in both the real and synthetic versions. If those numbers differ, you've found the first thing to fix before any model trains on it.

What are You Looking For?

esc