Sam Austin on October 8, 2026

YData Synthetic Tutorial: Open-Source Synthetic Data Library

YData Synthetic Tutorial: Open-Source Synthetic Data Library
Contents

Abstract library of generative models representing open-source synthetic data tooling

Figure 1: An open-source library being deprecated doesn't delete it — but it should change what you build on next

Another status update needed before diving in: ydata-synthetic, the open-source library this tutorial was requested for, has been deprecated by its own maintainers in favor of a commercial, API-based product called ydata-sdk. The GitHub repository and PyPI page both carry an explicit banner: "we are now ydata-sdk," and independent package health analysis confirms the original library hasn't seen a meaningful release in the better part of a year, with maintenance status flagged as inactive. Let's cover what actually happened, what still technically works, and what the genuinely maintained open-source alternative looks like today.

What ydata-synthetic Actually Was

Launched in 2020, ydata-synthetic was a pioneering open-source package built with an explicit educational mission: helping users understand generative models for synthetic data generation, rather than being purely a production tool. It covered both tabular data through a CTGAN-based conditional architecture and time-series data through a TimeGAN implementation, with a Streamlit-based UI layered on top for people who wanted a guided interface rather than writing Python directly. For several years, it was a genuinely popular starting point for anyone learning synthetic tabular data generation, with real community engagement reflected in over 1,600 GitHub stars.

What Changed

YData, the company behind the library, made a deliberate strategic shift: rather than continuing active development on the open-source package, they redirected users toward ydata-sdk, described as providing superior performance, precision, and ease of use by automatically selecting and optimizing the best generative model for your specific data, rather than requiring you to choose between CTGAN, TimeGAN, or other architectures yourself.

The practical consequence for anyone actually trying to use the old library: imports break. A documented GitHub issue shows exactly this — from ydata_synthetic.synthesizers.regular import RegularSynthesizer, the standard import shown in the library's own example notebooks, no longer works in version 2.0.0, replaced by a deprecation warning redirecting to ydata.sdk.synthesizers instead. If you follow an older tutorial's exact code, including possibly earlier versions of this very kind of article, you'll hit that exact breakage.

Worth noting too: PyPI metadata for the package shows it was renamed again, to fg-data-synthetic, with a release as recent as April 2026, though the broader package health signals — no meaningful development activity, maintenance flagged as inactive — suggest this is closer to a maintenance placeholder than active ongoing development. Treat any version you find with real caution and verify current behavior directly before building anything on it.

What ydata-sdk Actually Is

If you want to stay within YData's ecosystem, ydata-sdk is the actual current product:

pip install ydata-sdk

It positions itself as a broader data-centric toolkit, not just synthetic data generation, bundling connectors, metadata management, and data quality profiling alongside synthesis, with a single API that automatically selects a generative model for your data rather than requiring manual architecture choice. For a full UI-driven experience, YData also offers YData Fabric, which covers the whole pipeline from data preparation through synthetic data generation and evaluation, with a community version available as a free trial.

The honest framing here: this is a commercial product with a free tier entry point, not a pure open-source library the way the original ydata-synthetic was. If your organization is fine adopting a vendor relationship for this capability, it's worth evaluating directly on YData's own site rather than trying to reconstruct the deprecated open-source version's old API.

The Genuinely Maintained Open-Source Alternative

If what you actually want is a free, actively maintained, pure open-source library for tabular and time-series synthetic data generation, the more reliable choice right now is the Synthetic Data Vault (SDV) ecosystem, which includes the standalone CTGAN package used throughout this series' other synthetic data pieces — finance, retail, and manufacturing alike — as covered in the SDV walkthrough elsewhere in this series.

pip install sdv
from sdv.single_table import CTGANSynthesizer
from sdv.metadata import SingleTableMetadata
import pandas as pd

real_data = pd.read_csv("your_data.csv")

metadata = SingleTableMetadata()
metadata.detect_from_dataframe(real_data)

synthesizer = CTGANSynthesizer(metadata)
synthesizer.fit(real_data)

synthetic_data = synthesizer.sample(num_rows=10_000)

SDV covers single-table, multi-table (relational), and sequential (time-series) data generation under one actively maintained umbrella, with a genuinely large and current user base, and it doesn't carry the same deprecation risk ydata-synthetic currently does. For most of the use cases ydata-synthetic originally targeted — tabular synthesis, time-series synthesis — this is the more defensible choice to build real infrastructure around today.

If You Specifically Need the Old ydata-synthetic Behavior

If you've got existing code built against the pre-deprecation ydata-synthetic API and need it to keep working rather than migrating immediately, a few practical steps matter.

  • Pin your installed version explicitly rather than letting pip install ydata-synthetic pull whatever the latest tag happens to be, since the import paths and available classes have genuinely changed between versions, as the documented GitHub issue shows directly.
  • Check the deprecated documentation YData itself still hosts, explicitly linked from their current docs site as "older version with open-source models" — this is the maintainers' own acknowledgment that the old API still exists in documented form even though it's no longer the recommended path forward.
  • Budget real time for an eventual migration, either to SDV for a pure open-source path, or to ydata-sdk if you're willing to adopt the commercial product, since a library with inactive maintenance and no recent meaningful releases is not something to build new, long-term infrastructure on regardless of whether old code currently still runs.

A Practical Comparison for Your Decision

If you want a free, actively maintained, pure open-source library with no vendor relationship, SDV and CTGAN are the more defensible current choice, covering tabular, relational, and time-series data generation with ongoing development behind it.

If you want a managed, single-API experience that automatically picks the right generative model for you and you're fine with a vendor relationship, ydata-sdk is YData's own current direction, worth evaluating directly against your specific data types and budget.

If you've got legacy code on the original ydata-synthetic API and migrating isn't feasible immediately, pin your version explicitly, reference the maintainers' own archived documentation, and treat this as a bridge, not a destination, given the clear, maintainer-acknowledged deprecation.

If you need NVIDIA's enterprise-grade synthetic data tooling specifically, covered in the Gretel piece elsewhere in this series, that's yet another distinct path, now routed through NeMo's Data Designer and Safe Synthesizer microservices under NVIDIA AI Enterprise, with its own separate sales-gated access model.

Common Mistakes to Avoid

  • Following an older tutorial's exact import paths without checking current library behavior first is the single most common way people hit the documented RegularSynthesizer breakage directly — verify against the current state of whatever library you're actually installing before writing real code against it.
  • Building new, long-term infrastructure on a library explicitly flagged by its own maintainers as deprecated, with inactive maintenance confirmed independently, creates technical debt you'll need to address eventually regardless of whether it currently still runs.
  • Assuming "open source" and "actively maintained" are the same thing — ydata-synthetic's GPL-3.0 license and public GitHub repository remain technically true even as the actual development and support have clearly moved elsewhere.
  • Not checking whether a seemingly-current PyPI release (like the April 2026 fg-data-synthetic rename) reflects genuine active development versus a maintenance placeholder — package health signals like release cadence and issue activity tell a more reliable story than a recent-looking version number alone.
CoverBookDescriptionGet it
Cover of “Generative Deep Learning” Generative Deep Learningby David Foster the conceptual ground ydata-synthetic was built to teach: GANs, VAEs, and how generative models learn distributions, independent of which library's API you end up using. View on Amazon
Cover of “Designing Machine Learning Systems” Designing Machine Learning Systemsby Chip Huyen the dependency-discipline framing for this whole piece: evaluating tool maturity, maintenance risk, and data fitness before infrastructure gets built around a library. View on Amazon
Cover of “Machine Learning Engineering” Machine Learning Engineeringby Andriy Burkov practical guidance on productionizing models with maintained, well-understood components rather than convenient-but-stagnant ones. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

Is ydata-synthetic still maintained?

No. Its own maintainers deprecated it in favor of the commercial, API-based ydata-sdk: the GitHub repository and PyPI page both carry an explicit banner reading "we are now ydata-sdk," and independent package health analysis flags the original library's maintenance status as inactive, with no meaningful release in the better part of a year. The GPL-3.0 license and public repository remain technically true even as development has moved elsewhere.

What exactly broke in ydata-synthetic 2.0.0?

The standard import from ydata_synthetic.synthesizers.regular import RegularSynthesizer — the one shown in the library's own example notebooks and older tutorials — no longer works in version 2.0.0, replaced by a deprecation warning redirecting to ydata.sdk.synthesizers. PyPI metadata also shows a rename to fg-data-synthetic with an April 2026 release, though health signals suggest that's closer to a maintenance placeholder than active development.

What is the best maintained open-source alternative?

The Synthetic Data Vault (SDV) ecosystem, which includes the standalone CTGAN package used across this series' synthetic data pieces. SDV covers single-table, multi-table (relational), and sequential (time-series) generation under one actively maintained umbrella with a large current user base, and it doesn't carry ydata-synthetic's deprecation risk — making it the more defensible choice for new infrastructure.

I have existing code on the old ydata-synthetic API — what now?

Pin your installed version explicitly instead of letting pip pull the latest tag, since import paths and available classes changed between versions; reference the deprecated documentation YData still hosts, linked from their current docs as "older version with open-source models"; and budget real time for an eventual migration to SDV for an open-source path or ydata-sdk for the commercial one. Treat the pinned setup as a bridge, not a destination.

Wrapping Up

ydata-synthetic was a genuinely useful, pioneering open-source library for learning and using synthetic tabular and time-series data generation, but its own maintainers have explicitly redirected users toward the commercial ydata-sdk product, and independent package health signals confirm the original library's maintenance is now inactive. For anyone building new infrastructure today, SDV and CTGAN remain the more defensible, actively maintained open-source path for the same core tabular and time-series synthesis use cases, with no vendor relationship required.

If you landed here specifically wanting to follow a ydata-synthetic tutorial, the responsible next step is checking which of the three paths — SDV's open-source route, YData's own current commercial product, or migrating existing pinned-version code deliberately — actually fits your situation, rather than building against an API its own creators have already moved away from.

What are You Looking For?

esc