Sam Austin on October 8, 2026

Synthetic Data for Autonomous Vehicles: Simulating Driving Scenarios

Synthetic Data for Autonomous Vehicles: Simulating Driving Scenarios
Contents

Car driving on a mountain road at dusk representing rare driving scenarios

Figure 1: The scenario that matters most is the one your fleet will almost never encounter

A self-driving system needs to handle a child's ball rolling into the road at dusk in heavy rain, and no fleet will drive enough miles to see that scenario often enough to learn from it. Real data collection and annotation for safety-critical systems is slow and expensive, and the situations that matter most are the ones that almost never happen. That mismatch is exactly why autonomous vehicle (AV) teams lean on synthetic data harder than almost any other field.

The tooling has changed fast, though. "Driving simulator" now covers three quite different approaches, and knowing which one fits which job is the useful part.

Why the Long Tail Drives Everything

Normal driving is easy to collect and mostly uninteresting. The risk lives in the long tail: unusual road users, odd lighting, sensor glitches, aggressive cut-ins, construction zones, and combinations of them. A model trained on mostly ordinary miles can look excellent on average and still fail badly on rare events.

Synthetic data attacks this directly. You can generate rare hazards on demand, vary weather and lighting systematically, and get perfect labels for free — including exact object positions, depth, and segmentation — without paying humans to annotate each frame.

Approach One: Classic Game-Engine Simulators

The traditional route builds a virtual world from 3D assets and simulates cameras, lidar, and vehicle physics inside it. CARLA is the widely used open-source option, with controllable maps, weather, traffic, and sensor models. Scenarios are typically defined declaratively, often using standards like ASAM OpenSCENARIO, so you can script "pedestrian crosses 20 meters ahead while a vehicle merges" and replay it with variations:

# scenario: child's ball rolls into the road at dusk, heavy rain
scenario = {
    "name": "ball_dusk_rain",
    "ego_speed_kmh": 40,
    "trigger": {"object": "ball", "crosses": "ego_lane", "lead_time_s": 2.4},
    "actor": {"type": "child", "emerges_after_ball": 0.8},
    "weather": {"precipitation": 0.8, "sun_elevation_deg": -2},
}

Run it repeatedly with python sweep.py and you get a parameter sweep over timing, distance, and lighting from one definition.

The strengths are control and labels. You decide exactly what happens, you can sweep parameters like speed, distance, and timing, and you can reproduce a failure deterministically. The weakness is the visual gap: rasterized game-style rendering looks different from real camera footage, and perception models trained on it can stumble on real images. Treat older simulators mentioned in the literature, such as LGSVL and AirSim, with some care, since several have stopped being actively developed, and verify maintenance status before building on one.

Approach Two: Neural Reconstruction of Real Drives

The newer idea starts from real data. Reconstruction-based methods convert recorded driving logs into renderable neural scenes using techniques like Neural Radiance Fields (NeRF) or 3D Gaussian Splatting, so you can re-render the same street from new camera positions or replay it with changes. Some methods go further and model dynamic actors, such as pedestrians and vehicles, as separate Gaussian-splat objects in a scene graph, so you can move or remove them.

NVIDIA's Omniverse NuRec is a prominent implementation. It reconstructs real multi-sensor driving data into interactive 3D environments, integrates with CARLA and NVIDIA's AlpaSim, and targets regression testing and safety evaluation, with photoreal playback in under 25 milliseconds per frame. Variants aim to speed reconstruction up, generating a Gaussian scene in a single forward pass.

This approach narrows the visual gap, because the scene originates from real sensor data rather than hand-modeled assets. Its limitation is that you can only move so far from the recorded drive before the reconstruction runs out of information, such as areas the original cameras never saw.

Approach Three: Generative World Models

The third approach uses generative video models trained on driving data to produce new scenes, conditioned on controls like layout, depth, segmentation, or maps. NVIDIA's Cosmos platform is one example. Cosmos-Drive-Dreams is a synthetic data generation pipeline built to produce challenging scenarios for perception and driving policy training, with its toolkit, dataset, and model weights released openly. Related research systems such as GAIA-2 and DriveDreamer pursue similar goals from different angles — the same family of generative machinery covered in the diffusion model guide elsewhere in this series.

Cosmos Transfer illustrates how this fits with simulation: it can produce photorealistic video from ground-truth 3D simulation scenes or from spatial control inputs like depth, segmentation, edges, and HD maps. That combination is powerful, because the simulator supplies exact ground truth and structure while the generative model supplies realistic appearance. Companies in the AV toolchain space, like Foretellix, are integrating reconstruction, sensor simulation, and generative transfer to build physically grounded synthetic scenarios.

The tradeoff is fidelity to your conditioning. Generated video may add plausible but wrong details, so labels are only trustworthy for what you explicitly controlled, and subtle physical inconsistencies can slip through.

How the Pieces Fit Together

Increasingly, teams chain these approaches into a data flywheel. Real drives get reconstructed into interactive scenes. A scenario tool edits those scenes by adding a cut-in or a stalled vehicle, changing the weather, or shifting traffic. A generative model refines appearance across lighting and conditions. The result feeds perception and policy training, and failures found in testing become new scenarios. Curation tools help with the other end of the loop, retrieving rare real scenarios from large fleet archives so you know which gaps to fill synthetically.

Here's a simple way to think about which approach to use:

  • Need precise, repeatable scenario control (cut-in timing sweeps, regression tests)? Use a game-engine simulator like CARLA with declarative scenarios.
  • Need realistic appearance tied to real places? Use neural reconstruction of logged drives.
  • Need lots of visual variety (weather, lighting, rare conditions)? Use generative world models, ideally conditioned on simulator ground truth.
  • Need end-to-end evaluation of a driving policy? Use closed-loop simulation, where the vehicle's own decisions change what happens next.

Closing the Sim-to-Real Gap

The central problem is the same one that appears throughout this series: models trained on synthetic data can latch onto synthetic quirks. For rendering specifically, recent work in robotics simulation suggests the representation matters. A report on Gaussian-splat rendering for robot learning found that swapping mesh rasterization for splat-based rendering raised zero-shot sim-to-real success to roughly 86%, versus about 97.5% for models trained on real data, narrowing but not eliminating the gap. That's robotics, not driving, and the source is a secondary analysis, so treat it as directional evidence rather than an AV benchmark. It does match the broader pattern: rendering realism helps, and real data still sets the ceiling.

Mixing is the practical answer. In the Cosmos-Drive-Dreams experiments, the researchers trained with a fixed 50/50 ratio of synthetic to real data per epoch while varying the number of real clips from 2,000 to 20,000, which shows how teams study how much real data synthetic data can complement. Run similar ablations on your own task: vary the real data volume and the synthetic share, then measure on a held-out real evaluation set — the same discipline the synthetic data for computer vision piece recommends for image pipelines.

Open-Loop vs Closed-Loop Evaluation

Two evaluation styles behave very differently. Open-loop evaluation feeds recorded or synthetic frames to a model and scores its predictions, which is cheap but ignores consequences: the model's decisions don't change the next frame. Closed-loop evaluation lets the model drive inside a simulator so its decisions shape what it sees next, revealing compounding errors that open-loop tests miss. Closed-loop testing needs a simulator realistic enough that the model isn't exploiting simulator artifacts, which is where reconstruction and generative techniques earn their keep.

A Practical Workflow

Start by listing the scenarios and conditions your system must handle, drawn from real incident data and standard hazard categories. Build a scenario library in a declarative format, with parameters you can sweep. Use reconstruction or generative appearance transfer where visual realism matters, but keep simulator ground truth for labels. Train on a mix of real and synthetic data, tune the ratio empirically, and always evaluate on real held-out data. Track coverage: which scenario categories and parameter ranges your tests actually exercised. Then feed real-world failures back into the scenario library.

Common Pitfalls

  • Treating simulator success as safety evidence. Passing a large scenario suite shows coverage of what you modeled, not performance on what you didn't imagine, and safety cases need the simulator's own fidelity validated against reality.
  • Trusting generated video's labels beyond what you conditioned on. A world model that invents a pedestrian's pose or a sign's text has produced appearance, not ground truth.
  • Using only open-loop metrics, which can look fine while closed-loop behavior fails.
  • Over-weighting rare scenarios. If synthetic hazards dominate training, the model may become oversensitive and brake for phantom events, so balance rare cases against ordinary driving.
  • Ignoring sensor realism. Cameras, lidar, and radar each have noise, blur, and failure modes, and a simulator that renders idealized sensor output teaches the model an unrealistic world.
  • Locking into one vendor stack without checking whether scenarios, assets, and reconstructions can move between tools, since this field changes quickly.
CoverBookDescriptionGet it
Cover of “Probabilistic Robotics” Probabilistic Roboticsby Thrun, Burgard, and Fox the estimation foundation behind sensor models, localization, and why noisy simulated sensors behave differently from idealized ones. View on Amazon
Cover of “Robotics: Modelling, Planning and Control” Robotics: Modelling, Planning and Controlby Siciliano et al. vehicle and vehicle-dynamics modeling, scenario geometry, and the control side of closed-loop evaluation. View on Amazon
Cover of “Multiple View Geometry in Computer Vision” Multiple View Geometry in Computer Visionby Hartley and Zisserman the geometric machinery underneath NeRF and Gaussian-splat reconstruction of real drives: cameras, poses, and epipolar geometry. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

Why can't real driving miles cover rare AV scenarios?

Because the scenarios that matter live in the long tail: unusual road users, odd lighting, sensor glitches, aggressive cut-ins, construction zones, and combinations of them. Normal driving is easy to collect and mostly uninteresting, while a child's ball rolling into the road at dusk in heavy rain may never happen enough times in a fleet's lifetime to learn from. Synthetic data lets you generate those hazards on demand, vary weather and lighting systematically, and get perfect labels for free.

What's the difference between the three simulation approaches?

Game-engine simulators like CARLA build a virtual world from 3D assets and give you precise, repeatable scenario control and exact labels, at the cost of a visual gap from real footage. Neural reconstruction turns recorded drives into renderable scenes using NeRF or 3D Gaussian Splatting, narrowing the visual gap but only covering what the original sensors saw. Generative world models produce new video conditioned on controls like depth, segmentation, or maps, offering the most visual variety but with labels trustworthy only for what you explicitly conditioned on.

What is open-loop versus closed-loop evaluation?

Open-loop evaluation feeds recorded or synthetic frames to a model and scores its predictions; it's cheap but ignores consequences, because the model's decisions don't change the next frame. Closed-loop evaluation lets the model drive inside a simulator so its decisions shape what it sees next, revealing compounding errors that open-loop tests miss. Closed-loop testing requires a simulator realistic enough that the model isn't exploiting simulator artifacts.

Does synthetic data alone prove an AV is safe?

No, and responsible teams don't claim it can. Passing a large scenario suite shows coverage of what you modeled, not performance on what you didn't imagine, so safety cases need the simulator's own fidelity validated against reality. The realistic path is a rigorous scenario library with measured coverage, training on a mix of real and synthetic data, and keeping real held-out data as the final judge of performance.

Wrapping Up

Synthetic data for autonomous vehicles now spans three complementary approaches: game-engine simulators for precise, repeatable scenario control, neural reconstruction for realistic scenes grounded in real drives, and generative world models for visual variety and long-tail conditions. The strongest pipelines combine them, using simulator ground truth for labels and generative models for appearance, then mixing the result with real data and validating on real evaluation sets.

Can simulation alone prove a vehicle is safe? No, and responsible teams don't claim it can. But it's the only practical way to expose a system to the rare, dangerous situations that real fleets almost never encounter, which is why building a rigorous scenario library, measuring coverage, and keeping real data as the final judge is the realistic path.

What are You Looking For?

esc