Contents
Figure 1: The tool you bookmarked may not be the tool that exists — check before you build on it
Before writing a line of tutorial content, you need to know something that changes this entire request: Gretel's self-serve synthetic data platform is effectively discontinued. NVIDIA acquired Gretel in March 2025 for a reported $320 million-plus, and rather than continuing to operate it as an independent product, Gretel's technology and roughly 80-person team were folded directly into NVIDIA's own generative AI tooling. If you're looking for the Gretel API and SDK that used to have a self-serve free tier, that product, as it existed, isn't the thing to tutorial anymore. Let's cover honestly what happened and what you'd actually use today.
What Actually Happened to Gretel
Gretel launched in 2019 around a single clear promise: privacy-preserving synthetic data, accessible through an API and SDK with a genuine self-serve free tier, serving organizations in finance, healthcare, and the public sector who needed realistic data without the privacy exposure of using real records. That product existed for roughly six years before the acquisition.
Post-acquisition, Gretel's synthetic-data technology was absorbed to power NVIDIA's own AI platform rather than continuing as an independent, directly-accessible product. The underlying technology re-emerged inside NVIDIA NeMo as two distinct microservices: Data Designer, for synthetic dataset generation, and Safe Synthesizer, for privacy-preserving synthesis specifically. Both are distributed under NVIDIA AI Enterprise, which is sales-gated, meaning there's no public self-serve pricing or signup flow the way Gretel's original platform offered. If you want to use what used to be Gretel's technology today, you're going through an NVIDIA enterprise sales conversation, not a self-serve signup.
Why This Matters for Anyone Who Bookmarked a Gretel Tutorial
If you find an older tutorial walking through gretel-client pip installs and free-tier API keys, know that it's describing a product experience that's gone. The specific appeal that made Gretel popular for quick prototyping — sign up, get an API key, generate synthetic data in a notebook within minutes — no longer exists in that form. What NVIDIA kept is the underlying generative technology, repackaged for enterprise deployment rather than the developer-friendly, low-friction access point Gretel built its reputation on.
This is worth knowing before you invest time trying to follow stale documentation or forum posts describing a signup flow that may no longer function the way it's described.
What to Actually Use Instead
Given that reality, here's where to look depending on what you actually need.
If your organization already has, or is willing to pursue, an NVIDIA AI Enterprise relationship, NeMo Data Designer and Safe Synthesizer are the direct technical successors, carrying forward Gretel's actual generation and privacy-preservation technology, just wrapped in NVIDIA's enterprise packaging and sales process rather than a self-serve developer product. Worth evaluating specifically if you're already deep in NVIDIA's broader AI infrastructure stack, since this would sit naturally alongside tooling you may already be using.
If you want something you can actually sign up for and start using today without an enterprise sales conversation, MOSTLY AI is a genuinely comparable alternative worth evaluating first. It offers a hosted platform plus a Python SDK, and notably, its SDK itself is open source, even though the hosted platform is a commercial product. For model-training synthetic data specifically — tabular, time-series, the use cases Gretel originally targeted — MOSTLY AI occupies very similar ground with an access model closer to what Gretel used to offer.
If you want no vendor relationship at all, the open-source ecosystem covered elsewhere in this series remains fully available and unaffected by any of this: CTGAN and the broader Synthetic Data Vault (SDV) project for tabular data generation, and domain-specific tools like RetailSynth for retail-specific simulation or Defect-GAN and similar for manufacturing defect synthesis. None of these require any vendor relationship, and they cover a meaningful share of what Gretel's original platform did, with more hands-on setup required in exchange for no sales gate and no licensing cost.
A Basic Generation Example With an Open Alternative
Since a Gretel-specific walkthrough isn't the useful thing to hand you right now, here's the equivalent workflow using CTGAN directly — the same technique covered for finance, retail, and manufacturing synthetic data elsewhere in this series — since it's freely available and solves much of the same core problem Gretel's tabular synthesis did:
pip install ctgan
from ctgan import CTGAN
import pandas as pd
real_data = pd.read_csv("your_data.csv")
discrete_columns = ["category_column", "status_column"]
model = CTGAN(epochs=300)
model.fit(real_data, discrete_columns)
synthetic_data = model.sample(10_000)
synthetic_data.to_csv("synthetic_output.csv", index=False)
This gets you a working synthetic dataset without any vendor dependency at all. For privacy guarantees specifically, differential privacy mechanisms (covered in the financial synthetic data piece in this series) are worth layering on top if your use case has genuine privacy-sensitivity requirements — the kind Gretel's original platform built tunable, mathematically-proven privacy settings around directly.
What to Check Before You Commit to Any Path
Given how fast this specific landscape has shifted, verify a few things directly before building real infrastructure around any synthetic data tool right now:
- Confirm MOSTLY AI's current pricing and feature set directly on their site rather than trusting a secondhand comparison — commercial synthetic data platforms update pricing and tiers frequently.
- If you're considering the NVIDIA NeMo path, reach out through NVIDIA's enterprise channels directly to understand actual current capability and pricing, since sales-gated enterprise products don't publish the kind of self-serve documentation a tutorial could responsibly walk through anyway.
- If you're going the open-source route, check which specific library — SDV, CTGAN, or a domain-specific tool — actually fits your data type (tabular, time-series, text), since Gretel's original appeal was partly in supporting several data types under one unified platform, something you'll need to assemble yourself from separate open-source tools now.
Common Pitfalls
- Following a stale Gretel tutorial end to end —
gretel-clientinstalls and free-tier signup flows describe a product experience that no longer exists in that form, and time spent debugging a dead signup flow is time not spent on the actual task. - Assuming the NeMo path is a drop-in replacement because it's "the same technology" — the generation tech is, the access model isn't: sales-gated enterprise packaging is a fundamentally different workflow than a self-serve API key.
- Skipping privacy validation because a managed platform "handles privacy" — layer differential privacy and check the actual guarantees for your use case rather than inheriting a vendor's defaults without reading them.
- Picking tools per data type with no plan and ending up with four disconnected pipelines — decide up front whether you need tabular, time-series, and domain-specific generation, since assembling Gretel's unified-platform experience from separate open-source tools is the main real cost of the no-vendor route.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Synthetic Data Generation with Generative Models | survey-level coverage of the GAN and VAE approaches under tools like CTGAN, useful for understanding what any of these platforms are actually doing when they claim statistical fidelity. | View on Amazon |
![]() |
Designing Machine Learning Systems | the evaluation discipline that matters once you've generated data: whether it's fit for the downstream task, and how to validate that before training on it. | View on Amazon |
![]() |
Privacy-Preserving Data Mining | the differential privacy and membership-inference background that decides whether your synthetic dataset actually protects the real records it was learned from. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What happened to Gretel.ai's self-serve synthetic data platform?
NVIDIA acquired Gretel in March 2025 for a reported $320 million-plus, and rather than continuing it as an independent product, Gretel's technology and roughly 80-person team were folded into NVIDIA's own generative AI tooling. The self-serve signup, free tier, and API keys that made Gretel popular for quick prototyping no longer exist in that form.
Where did Gretel's technology end up?
Inside NVIDIA NeMo as two microservices: Data Designer for synthetic dataset generation and Safe Synthesizer for privacy-preserving synthesis. Both ship under NVIDIA AI Enterprise, which is sales-gated — no public self-serve pricing or signup flow the way Gretel's original platform offered, so using it today means an enterprise sales conversation rather than a pip install and an API key.
What should I use instead of Gretel today?
If you want a managed platform you can sign up for today, MOSTLY AI is the closest comparable option — hosted platform plus a Python SDK whose SDK itself is open source, covering the same tabular and time-series model-training use cases Gretel targeted. If you prefer no vendor relationship at all, CTGAN and the SDV project cover tabular generation, with domain tools like RetailSynth and Defect-GAN for retail and manufacturing, at the cost of more hands-on setup.
How do I generate synthetic tabular data without Gretel?
Install CTGAN, load your CSV, list the discrete columns, fit the model, and sample: roughly ten lines of Python that produce a synthetic dataset with no vendor dependency. For genuine privacy requirements, layer differential privacy mechanisms on top — the kind of tunable, mathematically-proven privacy settings Gretel's original platform wrapped around its own generators.
Wrapping Up
Gretel as a self-serve synthetic data platform is a product that existed, was genuinely well-regarded, and no longer operates the way it did before NVIDIA's March 2025 acquisition. Its technology lives on inside NVIDIA NeMo as Data Designer and Safe Synthesizer, but gated behind an enterprise sales process rather than the developer-friendly signup that made the original product popular for quick prototyping and smaller-scale projects.
If you came here looking for a quick-start Gretel tutorial, the honest answer is that the quick-start version of this tool doesn't exist anymore. MOSTLY AI is the closest thing to a direct, still-self-serve successor, and the open-source ecosystem — CTGAN, SDV, and the domain-specific tools covered elsewhere in this series — remains fully available with no vendor relationship required at all. Pick based on whether you need a managed platform's convenience or you're fine assembling the pipeline yourself.


