Contents
Figure 1: Three scores in one table — fidelity, utility, privacy
You spent three weeks generating synthetic data, and your model still flopped on real customers. The generator looked great on paper, but nobody had checked whether the data held up.
This guide fixes that. You'll learn how to measure fidelity and utility, how the popular tools compare, and how to run your own benchmark without losing a weekend.
Why Synthetic Data Benchmarks Matter
Every vendor claims their synthetic data is "realistic." Realistic according to whom? Without a benchmark, you take that claim on faith, and faith isn't a data strategy.
A good benchmark answers two questions. Fidelity asks whether the synthetic data looks like the real data. Utility asks whether the synthetic data does the job you need.
Those two questions sound identical, but they aren't. It's entirely possible to match every column distribution and still wreck a downstream model — keep reading and you'll see why.
Fidelity: Does the Fake Data Look Real?
Fidelity measures how closely synthetic data mirrors the statistical properties of the original. Think of it as a lie detector test for your generator — you want it to pass in the worst way.
Column-Level Checks
Start simple. Compare each column in the synthetic set against the real one.
- Histograms and density plots show whether numeric distributions match.
- Kolmogorov-Smirnov (KS) tests put a number on how far apart two numeric distributions sit.
- Category frequency comparisons reveal whether rare labels vanish or explode.
Column checks catch the obvious failures fast. They also lull you into a false sense of security, because a generator can nail every single column and still scramble how the columns relate to each other.
Relationship Checks
Here's where most generators sweat. Real data contains correlations — age and income, zip code and purchase history. Your synthetic data needs to keep them.
- Correlation matrix differences show whether pairwise relationships survive.
- Pairwise plots let you spot broken patterns by eye.
- Classifier two-sample tests train a model to tell real rows from fake ones. If the model can't tell them apart, you've got great fidelity.
The classifier test is worth dwelling on: it treats your data like a tiny Turing test, and it punishes lazy generators without mercy.
Utility: Does the Fake Data Actually Work?
Fidelity tells you the data looks right. Utility tells you the data works right. Skip this step and you'll ship a model built on a very convincing illusion.
Train on Synthetic, Test on Real (TSTR)
This is the gold standard for utility. Train a model on synthetic data, evaluate it on a held-out slice of real data, then compare against a model trained on real data.
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
real_model = LogisticRegression(max_iter=1000).fit(X_train_real, y_train_real)
syn_model = LogisticRegression(max_iter=1000).fit(X_train_syn, y_train_syn)
print("real baseline:", accuracy_score(y_test, real_model.predict(X_test)))
print("TSTR:", accuracy_score(y_test, syn_model.predict(X_test)))
A small gap means your synthetic data carries the signal your model needs. A big gap means your generator threw away something important. Always run TSTR with the same model type and the same hyperparameters on both sides, or the comparison means nothing — the TSTR discipline shows up throughout this series, including the NLP synthetic data article.
Query and Analytics Checks
Not everyone trains models. Plenty of teams use synthetic data for dashboards, testing, and demos. For them, utility means the numbers still tell the same story.
- Run the same aggregate queries on both datasets and compare results.
- Check that top-N lists (best customers, top products) keep their order.
- Confirm that time-based trends survive, including seasonality.
If your quarterly revenue chart shows a completely different shape, your synthetic data fails, no matter how pretty the histograms look.
Privacy: The Third Wheel
Fidelity and utility pull in one direction, and privacy pulls in the other. The closer your synthetic data hugs the real data, the higher the risk it leaks something about real people. A generator that simply copies your rows scores perfectly on fidelity and fails completely on privacy.
So measure it. Nearest-neighbor distance checks (distance-to-closest-record) flag synthetic rows that sit suspiciously close to real ones. Membership inference tests try to guess whether a specific record trained the generator. Treat both as mandatory, not optional — and note that elevated DCR scores alone can create a false sense of security, as the bias article warns.
FYI, privacy settings often cost you some utility. Expect that tradeoff and decide upfront how much you'll accept.
How the Popular Tools Stack Up
Results swing wildly depending on your data. A tool that shines on clean tabular data may struggle with messy multi-table databases. Use the notes below as a starting map, then verify everything on your own data.
Open-Source Options
- SDV (Synthetic Data Vault): Several models for single-table and multi-table data, plus built-in quality reports. It's the right first stop because it measures its own output — a huge plus for beginners.
- Synthcity: Targets researchers and includes a range of generators and evaluation metrics in one library. It rewards people who enjoy tinkering.
- CTGAN and TVAE (via SDV): Popular deep-learning generators for tabular data. They handle mixed column types well but need tuning and patience — and the TVAE writeup covers what the variational approach trades away.
Commercial Platforms
- Gretel: A managed service with generation and quality scoring built in. It suits teams that want less infrastructure work.
- MOSTLY AI: Focuses on privacy-safe tabular and multi-table data with built-in accuracy reports. It appeals to enterprises with strict compliance needs.
- YData: Combines profiling with synthetic generation, so you can inspect your data quality before and after.
Which one wins depends on your data, your budget, and your tolerance for configuration. Open-source tools give you control and cost you time. Commercial tools save time and cost you money.
How to Run Your Own Benchmark
Don't trust anybody's leaderboard. Build a small benchmark in a day or two, and you'll learn more than any blog post can teach you.
- Pick a representative dataset. Use real data with the quirks you care about: missing values, rare categories, skewed numbers.
- Split before you generate. Hold out a real test set that your generators never see.
- Generate with each tool. Use default settings first, then tune one change at a time.
- Score fidelity. Run column checks, correlation checks, and a classifier two-sample test.
- Score utility. Run TSTR on your actual task and compare against a real-data baseline.
- Score privacy. Run nearest-neighbor and membership inference checks.
- Compare everything in one table. Seeing all three scores side by side exposes the tradeoffs fast.
Track runtime and cost too. A generator that needs twelve hours for a modest dataset might not fit your workflow, however good its scores look — the same discipline the small datasets ladder applies to technique choice applies here.
Common Benchmarking Mistakes
Each one quietly ruins your results:
- Testing on the training data. Always evaluate on held-out real data.
- Trusting a single metric. One score never captures the whole picture, so use several.
- Ignoring rare categories. Generators love to drop the weird stuff, and the weird stuff often matters most.
- Skipping the baseline. Without a real-data comparison, your utility numbers float in a vacuum.
- Tuning one tool and leaving the rest on defaults. That rigs the contest.
Another trap: chasing a perfect score. Perfect fidelity usually signals memorization, and memorization defeats the whole point of synthetic data — the exact failure mode the fundamentals article flags as both a quality and a privacy problem.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Synthetic Data Generation: Methods and Applications | the landscape behind the tool list, with evaluation chapters you can map directly onto the three scores. | View on Amazon |
![]() |
Designing Machine Learning Systems | data-centric discipline: baselines, offline/online evaluation, and why the pipeline beats the model. | View on Amazon |
![]() |
Statistics Done Wrong | the measurement mistakes in the mistakes section, explained with the rigor benchmarking requires. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is the difference between fidelity and utility in synthetic data?
Fidelity asks whether the synthetic data looks like the real data — matching column distributions, correlations, and joint structure. Utility asks whether the synthetic data does the job you need — whether a model trained on it performs on real held-out data, or whether your dashboard queries return the same story. They can disagree: a dataset can match every column distribution and still wreck a downstream model, usually because relationships between columns didn't survive generation.
What is TSTR and why is it the standard utility check?
TSTR means train on synthetic, test on real: train a model on generated data, evaluate it on a held-out slice of real data, and compare against a model trained on real data. A small gap means your synthetic data carries the signal your model needs; a big gap means the generator threw away something important. Always run both sides with the same model type and the same hyperparameters, or the comparison means nothing.
How do you test privacy in synthetic data?
Two checks are mandatory rather than optional. Nearest-neighbor distance checks (distance-to-closest-record) flag synthetic rows that sit suspiciously close to real ones, and membership inference tests try to guess whether a specific record trained the generator. A generator that simply copies rows scores perfectly on fidelity and fails privacy, and privacy settings typically cost some utility — decide upfront how much of a tradeoff you'll accept.
How do I run a synthetic data benchmark myself?
Pick a representative real dataset with the quirks you care about, split and hold out a real test set before any generator sees data, generate with each tool using defaults first, then score fidelity (column checks, correlation checks, classifier two-sample test), utility (TSTR on your actual task versus a real-data baseline), and privacy (nearest-neighbor and membership inference). Compare all three in one table with runtime and cost, and tune every tool the same amount or you rig the contest.
Wrapping This Up
Benchmarking synthetic data comes down to three questions: does it look real (fidelity), does it work (utility), does it stay safe (privacy)? Answer all three, and you'll choose tools with confidence instead of crossed fingers.
Remember that no single tool wins everywhere. Open-source libraries like SDV give you flexibility, and commercial platforms like Gretel and MOSTLY AI save you effort. Your data decides the winner, not the marketing page.
So here's the homework: grab a small dataset this week, run two generators, and score them with TSTR. You'll spend an afternoon and gain months of clarity. Come back and tell me which one surprised you. :)


