Contents
Figure 1: An autoencoder squeezes data through a narrow latent space and rebuilds it — the VAE version just makes that space something you can actually sample from
CTGAN gets most of the attention in synthetic tabular data conversations, but the paper that introduced it also introduced a second model: TVAE, a variational autoencoder built for the exact same problem. In direct empirical comparisons, TVAE frequently outperforms CTGAN on fidelity, and it's the quieter, less-hyped half of what's become the standard tabular synthesis baseline pair. Let's cover how it actually works, where it beats GANs, and where it genuinely doesn't.
What a VAE Actually Does, Conceptually
A variational autoencoder learns two things at once: an encoder that compresses real data into a lower-dimensional latent representation, and a decoder that reconstructs data back out of that latent space. Unlike a plain autoencoder, a VAE doesn't learn a single fixed point in latent space for each input, it learns a probability distribution over latent representations, and training optimizes something called the Evidence Lower Bound (ELBO), a loss function balancing two competing goals: reconstruction accuracy, how well the decoder rebuilds the original input from its latent encoding, against a regularization term (KL divergence) that pushes the learned latent distribution toward a well-behaved, samplable shape, typically close to a standard normal distribution.
That second term is the entire reason a VAE is useful for synthetic data generation at all. Once training converges, you can sample random points directly from that well-behaved latent distribution — points the model never saw during training — and run them through the decoder to produce entirely new, synthetic records that follow the same statistical patterns as your real data.
TVAE: The Tabular Adaptation
TVAE adapts this general VAE framework specifically for mixed-type tabular data, introduced in the same 2019 paper that introduced CTGAN, and it shares several of CTGAN's preprocessing techniques directly. Both models encode categorical variables through similar procedures, and both handle continuous variables through mode-specific normalization, using a Gaussian Mixture Model to represent continuous columns that have multiple distinct modes — a salary column with distinct clusters around different pay bands, for instance, rather than assuming a single smooth bell curve the way naive normalization would.
Where TVAE actually differs from CTGAN is architectural, not just a difference in encoding tricks. CTGAN trains a generator and discriminator adversarially, with the generator never directly seeing real training records, only the discriminator's feedback about whether its output looks real. TVAE instead directly encodes real records into latent space and decodes them back, optimizing the ELBO loss straightforwardly, without any adversarial component or discriminator network at all.
from sdv.single_table import TVAESynthesizer
from sdv.metadata import SingleTableMetadata
import pandas as pd
real_data = pd.read_csv("your_data.csv")
metadata = SingleTableMetadata()
metadata.detect_from_dataframe(real_data)
synthesizer = TVAESynthesizer(metadata)
synthesizer.fit(real_data)
synthetic_data = synthesizer.sample(num_rows=10_000)
Notice this is nearly identical to the CTGAN and GaussianCopula workflows covered elsewhere in this series: SDV deliberately keeps a consistent API across its synthesizers specifically so swapping between them for comparison is close to a one-line change, as the SDV walkthrough shows directly.
Why TVAE Often Wins on Fidelity
This is worth stating plainly because it runs against what a lot of people assume: research directly comparing these methods has repeatedly found that TVAE outperforms CTGAN and other GAN-based models on fidelity in extensive empirical comparisons. One study specifically built on this observation, explicitly naming TVAE as "a frontrunner" to benchmark newer models against, precisely because its direct, non-adversarial training objective tends to capture a dataset's actual statistical structure more faithfully than adversarial training does.
The likely mechanism behind this: GAN training is notoriously unstable, prone to mode collapse, where the generator learns to produce only a narrow subset of the real data's actual diversity because that's what currently fools the discriminator, rather than genuinely covering the full distribution. A VAE's direct reconstruction objective doesn't have an adversary to fool, it's optimizing a clear, well-understood loss function throughout training, which tends to produce more stable, broadly faithful results even if the theoretical ceiling on fidelity for either architecture is similar.
Where CTGAN Still Has a Genuine Edge
The honest counterpoint, directly noted in research comparing the two architectures: CTGAN achieves differential privacy more easily than TVAE, specifically because its generator never directly sees the original training data, only indirect feedback through the discriminator. That architectural separation gives you a cleaner surface to apply formal privacy mechanisms against, since the component actually producing synthetic data isn't the one with direct access to real records.
TVAE's encoder, by contrast, directly processes real data as part of normal training, which means applying a rigorous differential privacy guarantee requires protecting that encoding step specifically, and research notes that existing DP-TVAE implementations retain features of the original algorithm that represent privacy leakage outside what gets formally accounted for in the stated privacy guarantee. If a provable, tight differential privacy bound is a hard requirement, this is a real, documented reason to lean toward CTGAN's architecture, or toward the copula-based approach with its own more mature DP literature, rather than assuming TVAE's fidelity advantage settles the question on its own.
Where This Is All Heading: VAEs Inside Diffusion Models
Here's a genuinely important current development worth knowing about: rather than being displaced by newer diffusion-based tabular generators, VAEs have become a foundational component inside them. TabSyn, described directly in recent research as a state-of-the-art approach for generating high-quality synthetic tabular data, works by first using a VAE architecture to transform raw tabular data into a continuous latent space, one that captures both inter-column dependencies and token-level representations, and only then running a diffusion process inside that VAE-constructed latent space, rather than applying diffusion directly to the raw, messy mix of numerical and categorical columns.
This reframes what a "VAE versus GAN versus diffusion" comparison even means for tabular data in 2026. The VAE isn't necessarily competing with diffusion models as an alternative generation technique, in architectures like TabSyn it's the mechanism that makes diffusion work well on mixed-type tabular data in the first place, by giving the diffusion process a clean, unified continuous space to operate in rather than forcing it to handle categorical columns directly. If you're evaluating the current frontier of tabular synthesis quality, TabSyn and similar VAE-plus-diffusion hybrids, not a standalone TVAE, represent where fidelity benchmarks currently top out.
A Direct Empirical Comparison Worth Internalizing
A recent financial-synthetic-data study ran four representative methods side by side: Gaussian Copula, CTGAN, TVAE, and TabDiff, a diffusion-based approach, specifically to measure quality, utility, and privacy tradeoffs together rather than any single metric in isolation — the same validation discipline the finance piece in this series argues for. No single architecture, copula, GAN, VAE, or diffusion, wins uniformly across fidelity, privacy, and computational cost simultaneously, and the honest answer for any real project is to benchmark a couple of candidates directly against your actual data and your actual downstream task, rather than picking one architecture based on reputation or whichever one you've heard about most.
Improvements on Vanilla TVAE
Worth knowing that TVAE itself isn't the final word even within the pure-VAE family. Research into Oblivious Variational Autoencoders (OVAE) found this refined architecture outperforming TVAE on real-world classification datasets in aggregate, with OVAE described as the best-performing system across the specific benchmark studied. "VAE for tabular data" is an active, still-evolving research area, not a solved problem where TVAE represents a permanent ceiling, worth periodically checking current literature if fidelity on your specific data type is a serious ongoing concern rather than a one-time project decision.
Practical Guidance for Choosing
Reach for TVAE as a genuinely strong first choice when fidelity to your real data's statistical structure matters more than a hard, formal privacy guarantee, and you want something more stable and less fussy to train than an adversarial GAN setup, without giving up the ability to model genuinely complex mixed-type data the way a simpler Gaussian copula might struggle with.
Reach for CTGAN specifically, even knowing TVAE often edges it out on raw fidelity, when a provable differential privacy bound is a hard requirement, since its architecture — the generator never directly touching real data — gives privacy mechanisms a cleaner surface to work against.
Reach for TabSyn or another VAE-plus-diffusion hybrid when you're chasing the current fidelity frontier specifically and are willing to accept more implementation complexity and training cost than either TVAE or CTGAN demands on their own.
Reach for a Gaussian or vine copula instead of any neural approach when interpretability, fitting speed, or a mature differential privacy literature matter more to you than squeezing out the last few points of fidelity a deep learning method might offer — the tradeoff is covered in the copulas piece.
Run more than one of these against your actual data before committing. The direct comparative research consistently finds that relative performance shifts meaningfully depending on dataset characteristics, and reputation alone — "TVAE beats CTGAN in general" — doesn't guarantee that holds for your specific columns and distributions.
Common Pitfalls
- Assuming CTGAN is the default, stronger choice simply because it gets more attention and tutorial coverage — direct empirical research consistently shows TVAE matching or beating it on fidelity, so test both rather than defaulting to whichever one shows up first in search results.
- Applying a DP mechanism to TVAE and assuming it carries the same clean privacy guarantee CTGAN's architecture more naturally supports — documented research flags that existing DP-TVAE implementations retain privacy leakage beyond what gets formally accounted for, verify this carefully rather than assuming parity between the two architectures' privacy properties.
- Treating TVAE as the ceiling of what VAE-based approaches can achieve, when both OVAE and VAE-plus-diffusion hybrids like TabSyn have demonstrated meaningful improvements over vanilla TVAE on real benchmarks.
- Picking one architecture based on general reputation rather than running your own comparison against your actual dataset — the research consistently shows relative performance shifting by dataset, so a one-size-fits-all architecture choice skips validation that genuinely matters.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
| Variational Autoencoders for Deep Learning | the ELBO, reparameterization trick, and latent-space sampling that TVAE's whole design rests on, explained properly rather than hand-waved. | View on Amazon | |
![]() |
Generative Deep Learning | VAEs, GANs, and diffusion models side by side, exactly the framing needed to understand why TabSyn combines the first and third rather than choosing between them. | View on Amazon |
![]() |
Designing Machine Learning Systems | the benchmarking discipline this piece keeps returning to: evaluate candidates on your own data and downstream task instead of trusting architecture reputations. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
Does TVAE really outperform CTGAN on fidelity?
Extensive direct empirical comparisons say yes more often than not. Research explicitly naming TVAE a frontrunner found its direct, non-adversarial reconstruction objective captures a dataset's actual statistical structure more faithfully than adversarial training, with the likely mechanism being GAN training's mode collapse risk. Run both on your own data anyway, since relative performance shifts with dataset characteristics.
When should I pick CTGAN over TVAE?
When a provable differential privacy bound is a hard requirement. CTGAN's generator never directly sees real training records, only discriminator feedback, which gives privacy mechanisms a cleaner surface to apply against. Research also notes existing DP-TVAE implementations retain features representing privacy leakage outside what the stated guarantee formally accounts for.
What is TabSyn and does it make TVAE obsolete?
TabSyn is a state-of-the-art tabular generator that uses a VAE to build a continuous latent space capturing inter-column and token-level structure, then runs diffusion inside that latent space instead of on raw columns. VAEs aren't competing with diffusion, they're what makes diffusion work well on mixed-type data, so standalone TVAE sits below the current fidelity frontier.
Is TVAE still the best pure-VAE option?
Not necessarily. Research into Oblivious Variational Autoencoders (OVAE) found the refined architecture outperforming TVAE in aggregate on real-world classification benchmarks, with OVAE described as best-performing in the study. VAE for tabular data remains an active research area rather than a solved problem where TVAE represents a permanent ceiling.
Wrapping Up
TVAE offers a genuinely strong, often underrated alternative to GAN-based tabular synthesis: a non-adversarial, directly-optimized reconstruction objective that frequently produces higher-fidelity synthetic data than CTGAN in head-to-head research comparisons, at the cost of a less straightforward path to formal differential privacy guarantees. The field has also moved past treating VAEs and diffusion models as competitors, modern state-of-the-art approaches like TabSyn use a VAE specifically to construct the clean latent space a diffusion process needs to handle mixed-type tabular data well.
Will TVAE always beat CTGAN, or always lose to a copula model, or always trail a diffusion-based approach? None of the above. Consistently, the honest answer from direct comparative research is that it depends on your specific data, and the only way to know for certain is running a couple of these candidates against your own dataset and your own downstream validation task before committing to one as your production pipeline.

