Sam Austin on October 8, 2026

Synthetic Data for Rare Event Detection: Fraud, Defects, and Anomalies

Synthetic Data for Rare Event Detection: Fraud, Defects, and Anomalies
Contents

Credit cards and payment terminal representing rare event detection in fraud data

Figure 1: 99.8% accuracy, almost no fraud caught — the rare event problem in one headline number

Your fraud model reports 99.8% accuracy and catches almost none of the actual fraud. That's not a bug, it's what happens when 99.8% of your data is legitimate and a model learns that always predicting "legitimate" is the safest bet. Rare event detection — fraud, equipment failures, manufacturing defects, security intrusions — is fundamentally a data scarcity problem wearing a class imbalance costume, and synthetic data is one of several tools for addressing it. It's also a tool with real, documented ways of making things worse if you use it carelessly.

This piece pulls together lessons from the finance, manufacturing, and retail synthetic data work covered elsewhere in this series, focused specifically on the shared problem underneath all of them.

Start With the Boring Fixes First

Before generating a single synthetic record, it's worth being honest that synthetic oversampling isn't always the right first move. Class weighting, adjusting your loss function to penalize missed minority-class examples more heavily, and decision threshold tuning, moving the cutoff for what counts as a positive prediction, address imbalance without touching your training data at all, and they carry none of the risks synthetic generation introduces. Always establish a baseline using these simpler techniques first, so you actually know whether synthetic data is adding value or just adding complexity.

Also change how you measure success before changing anything else. Accuracy is nearly useless for rare events — a model that never flags anything can score above 99%. Use precision, recall, F1, and especially precision-recall AUC, which stays honest about minority-class performance in a way ROC AUC can obscure when positives are extremely rare.

The Oversampling Ladder: SMOTE to Generative Models

SMOTE (Synthetic Minority Over-sampling Technique) is the standard starting point, and it works differently than the generative methods covered elsewhere in this series. Rather than learning a data distribution, SMOTE interpolates directly between existing minority-class samples — for two fraud instances, it generates a synthetic point somewhere along the line connecting them. It's simple, fast, and a strong baseline that's genuinely hard to beat in some settings. The practical entry point is one package install, pip install imbalanced-learn, and a few lines of code.

Its limitations are well documented, though. SMOTE tends to produce overly smooth or redundant samples, essentially connecting the dots between real fraud clusters, which can generate synthetic points too similar to the originals to add much new information. It can also overgeneralize, introducing synthetic points in inappropriate regions of feature space, and variants like Borderline-SMOTE carry a specific risk of pulling in noise from majority-class samples near the decision boundary. For mixed categorical and continuous data, research has gone further, with one study proving that SMOTE-NC, the most widely used variant for mixed feature types, is not coherent and does not preserve the underlying data structure in a mathematically precise sense.

GANs and VAEs learn the underlying distribution rather than interpolating between existing points, which in principle lets them capture complex, non-linear structure SMOTE's straight-line interpolation can't. Hybrid approaches combining both, like an ensemble framework pairing SMOTE's initial balancing with a Wasserstein GAN refining the result, exist specifically to capture the strengths of each. One credit card fraud study found a VAE-GAN pipeline could outperform SMOTE, which still remained a strong baseline throughout.

The Honest Evidence: No Universal Winner

Here's the part vendor and tutorial content tends to skip: the research doesn't consistently favor generative models over SMOTE. One intrusion detection study found SMOTE variants significantly enhanced minority class detection, especially for weaker classifiers, while GAN-generated synthetic data had little impact on classifier accuracy in that context. Another study found models trained on SMOTE-balanced data produced better results than those trained on ADASYN or GAN-generated data. Meanwhile, other studies, particularly on highly non-linear problems like credit card fraud, find GAN-based augmentation reduces majority-class bias and improves F1 and AUC over SMOTE.

There's also a consistent trade-off pattern worth knowing: SMOTE often boosts recall at the expense of precision, catching more real positives while also generating more false alarms — a trade-off that matters enormously when each false positive costs an analyst's time or a customer's blocked transaction. The practical conclusion is the same one that keeps surfacing across this series: relative performance depends on your specific data, so benchmark a couple of approaches against your actual problem rather than defaulting to whichever sounds more sophisticated.

The Most Important Rule: Synthetic Data Goes in Training Only

If you take one practice from this piece, take this one. Synthetic samples should only ever be added to your training set, never your validation or test sets. Evaluation data must contain only real examples, emulating what actually happens in deployment, where your model will face real transactions, real defects, and real anomalies, not synthetic ones generated by the same process that helped train it.

The subtler version of this mistake happens during cross-validation. If you oversample your entire dataset before splitting into folds, synthetic samples derived from a given real example can end up in your training fold while that real example itself sits in the validation fold — a form of leakage that inflates your performance estimates and produces a model that looks great in evaluation and disappoints in production. The correct pattern applies oversampling inside each training fold only:

from imblearn.pipeline import Pipeline
from imblearn.over_sampling import SMOTE
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score, StratifiedKFold

pipeline = Pipeline([
    ("smote", SMOTE(random_state=42)),
    ("clf", RandomForestClassifier(random_state=42)),
])

cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(pipeline, X, y, cv=cv, scoring="average_precision")

Using imbalanced-learn's own Pipeline rather than scikit-learn's ensures SMOTE gets applied only to each fold's training portion, never leaking synthetic samples derived from validation data into training. Run it with python train.py and compare scores against your class-weighted baseline before deciding synthetic data earned its place.

Fraud: Fidelity Isn't Enough

Fraud carries a specific lesson covered in depth in the financial synthetic data piece in this series: standard fidelity metrics can pass while behavioral signals that matter for detection remain meaningfully off. Fraud is a behavioral problem built on temporal bursts, velocity patterns, and shared-infrastructure signals, structure that standard tabular generators aren't explicitly built to preserve. The documented case: a velocity rule firing at a materially lower rate in synthetic fraud than in real fraud, meaning a threshold tuned against the synthetic data would be too permissive against real attacks.

Applied to oversampling, this means interpolating between fraud rows, as SMOTE does, can produce points that look plausible feature by feature while breaking the time-dependent relationships between transactions that actually characterize fraud. If your fraud features are engineered from sequences — velocity counts, time-since-last-transaction, behavioral aggregates — validate that your synthetic examples preserve those relationships specifically, not just the marginal distributions of individual columns.

Defects: Combine Real and Synthetic, Or Skip Defects Entirely

Manufacturing defect detection, covered in the manufacturing synthetic data piece in this series, has two genuinely distinct paths. One is generating synthetic defect images, using techniques like Defect-GAN or 3D-rendering pipelines, and combining them with real defect images, since the research consensus is that a mix of real and photo-realistic synthetic images outperforms either alone. The other is skipping defect examples altogether through anomaly detection, training only on images of good parts and flagging anything that deviates from that learned baseline.

The anomaly detection path deserves emphasis here precisely because it sidesteps the rare-event data problem entirely rather than trying to solve it. If you have zero or near-zero defect examples, as with a brand-new product line, you can't meaningfully oversample a minority class that essentially doesn't exist yet, and training on normal data alone becomes the only practical starting point. It also catches defect categories nobody anticipated, since it isn't matching against a predefined defect taxonomy.

Anomalies in Sensor and Network Data

Industrial IoT and security contexts face the same imbalance: equipment failures and intrusions are rare, and the research here mostly follows the same ladder — SMOTE-family methods as a strong baseline, generative models and hybrids like ensemble WGAN frameworks for cases where SMOTE's linear interpolation misses complex structure, and physics-based simulation, covered in the manufacturing piece, for failure modes too costly or dangerous to induce deliberately. Optimized sampling strategies that ensure synthetic data meaningfully extends the classifier's decision boundaries, rather than just adding redundant points near existing examples, represent the direction this research is heading.

A Practical Decision Path

Start with class weights and threshold tuning, and measure everything with precision-recall metrics rather than accuracy. If that baseline isn't good enough, try SMOTE or a variant suited to your data type, using the pipeline pattern above so oversampling stays inside training folds. If your data has complex non-linear structure and SMOTE's results look unconvincing, try a generative approach like CTGAN, TVAE, or a hybrid, benchmarking against your SMOTE baseline rather than assuming improvement. For image-based defect detection, combine real and synthetic images, and consider anomaly detection if real defect examples are scarce or absent. For behavioral domains like fraud, validate temporal and relational signal preservation specifically, not just column-level fidelity. Throughout, keep evaluation data real and untouched by any synthetic generation process.

Common Pitfalls

  • Oversampling before splitting data into train and test sets or cross-validation folds leaks synthetic samples derived from evaluation data into training, producing inflated performance estimates that don't survive contact with production.
  • Reporting accuracy as your primary metric for a rare event problem hides total failure on the minority class behind a headline number that looks excellent.
  • Assuming a GAN or VAE automatically beats SMOTE, when the research is genuinely mixed and several studies find SMOTE variants outperforming generative approaches on specific datasets.
  • Evaluating synthetic fraud or anomaly data purely on distributional fidelity without checking whether behavioral and temporal signals that matter for detection survived generation.
  • Ignoring the precision cost of oversampling, since boosting recall by generating more minority-class examples often increases false positives, a trade-off that carries real operational cost in fraud review queues and manufacturing line stoppages.
CoverBookDescriptionGet it
Cover of “Imbalanced Learning: Foundations, Algorithms, and Applications” Imbalanced Learning: Foundations, Algorithms, and Applicationsby Haibo He et al. the academic foundation for everything in the oversampling ladder: cost-sensitive learning, SMOTE family variants, and why the boring fixes come first. View on Amazon
Cover of “Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow” Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlowby Aurélien Géron the practical model-evaluation chapters that make precision-recall trade-offs, threshold tuning, and train/test discipline concrete. View on Amazon
Cover of “Introduction to Anomaly Detection” Introduction to Anomaly Detection the no-labelled-examples path: training on normal data only, which is the answer when the rare event hasn't happened enough times to oversample at all. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

Is SMOTE still worth using, or should I jump straight to a GAN?

SMOTE is still worth using first. The research evidence is genuinely mixed rather than stacked in favor of generative models: several intrusion detection and tabular studies find SMOTE variants matching or beating GAN-generated data, while other studies on highly non-linear problems like credit card fraud find GAN-based augmentation ahead. Benchmark a SMOTE baseline against any generative approach on your own data instead of assuming the more sophisticated method wins.

Why is accuracy useless for rare event detection?

Because a model that never flags anything can still score above 99% when the positive class is that rare, which is exactly what happens with a fraud model that catches no fraud. Accuracy measures overall correctness and lets total minority-class failure hide behind an excellent headline number. Use precision, recall, F1, and especially precision-recall AUC, which stays honest about minority-class performance in a way ROC AUC can obscure when positives are extremely rare.

Can synthetic samples go into my validation or test set?

No. Synthetic samples should only ever be added to your training set. Evaluation data must contain only real examples, because deployment means facing real transactions, real defects, and real anomalies, not synthetic ones produced by the same process that helped train the model. The subtler version of this mistake is oversampling the full dataset before splitting into cross-validation folds, which leaks synthetic samples derived from validation data into training and inflates your estimates.

What if I have zero or near-zero examples of the rare event?

If the minority class essentially doesn't exist yet, as with a brand-new product line, you can't meaningfully oversample it. Start with anomaly detection trained only on normal examples — good parts, legitimate transactions — and flag anything that deviates from that learned baseline, which also catches categories nobody anticipated. Class weighting and threshold tuning still apply, and synthetic generation becomes relevant once a small real seed set exists to anchor it.

Wrapping Up

Synthetic data for rare event detection works best as one tool in a disciplined workflow, not a first resort. Establish a baseline with class weighting and threshold tuning, measure with precision-recall metrics instead of accuracy, treat SMOTE as a strong and simple baseline that generative models don't consistently beat, and keep synthetic data strictly inside training folds with real, untouched data for evaluation. For fraud, validate behavioral signal preservation beyond column-level fidelity; for defects, combine real and synthetic images or consider anomaly detection when real examples are scarce.

Will synthetic oversampling solve your rare event problem on its own? Rarely, and the mixed research evidence says to expect improvement on some datasets and little or none on others. But with a proper baseline, honest metrics, and leak-free evaluation, you'll know within an afternoon whether it's actually helping your specific problem, which beats discovering after deployment that your 99.8% accuracy was never measuring what you needed it to.

What are You Looking For?

esc