Sam Austin on October 8, 2026

Procedural Scene Generation for Computer Vision Training

Procedural Scene Generation for Computer Vision Training
Contents

Forest scene rendered from procedural rules representing generated 3D training environments

Figure 1: Every tree, rock, and shadow generated from rules — no asset library involved

Recall the previous article's tool comparison of Unity Perception, Isaac Sim, and BlenderProc — all of them assumed something that article didn't dwell on: someone, somewhere, had to build the 3D scenes and objects those tools randomize. Domain randomization varies lighting, textures, and object placement within a scene, but the underlying assets — the sofa, the room layout, the tree — still came from a modeling pipeline, a purchased asset library, or a team of 3D artists. Procedural generation attacks that bottleneck directly: instead of authoring or sourcing 3D content, you write rules that generate it, producing unlimited variety from code rather than from a finite library.

Here's the concrete distinction worth building this article around: Infinigen, the Princeton system leading this space, is described directly as "not a collection of assets or synthetic images; instead, it is a generator" — every shape and material is created from scratch using mathematical rules, with no pre-made assets involved at all. And you've already seen this technique without it being named: the quadruped locomotion article's curriculum-based terrain progression — ramps, random-height boxes, bumpy ground — was procedural terrain generation, and Minecraft itself is a procedurally generated world. This technique is far more embedded across this series than a standalone "procedural generation" label suggested.

By the end of this guide, you'll understand how procedural generation differs from both hand-authored simulation and generative models like the GANs and diffusion systems covered elsewhere in this series, see how Infinigen and ProcTHOR actually differ, and understand an honest limitation current research is explicit about: "procedural systems tend to produce sterile, homogeneous layouts" — from a September 2026 paper — is genuinely the sentence that best captures where this field's honest frontier currently sits.

Three Different Ways to Get Synthetic Training Scenes

Worth being precise about this three-way distinction, since it's genuinely easy to conflate procedural generation with its two neighbors.

Hand-authored simulation — a human builds or sources 3D assets, then a tool like Unity Perception randomizes lighting, camera, and placement. Rich ground truth, but diversity is capped by the asset library you assembled.

Generative models (GANs, diffusion) — a neural network learned from real images produces new images, but labels come from the model's inference rather than from known geometry, and memorization or quality-validation concerns apply.

Procedural generation — mathematical rules and parameters create the 3D geometry, materials, and scene composition from scratch. Because the system created every object, it knows its exact geometry, giving the same perfect ground truth simulation provides, while the diversity comes from the rule space rather than a finite asset library.

Procedural systems construct large scene collections from rules and reusable asset libraries, and owing to their controllability, scalability, and repeatability, they're well suited for constructing synthetic datasets with automatically generated geometry, materials, and annotations — this is genuinely the combination that makes the technique distinct: the perfect labels of simulation plus diversity that scales without proportional human effort.

Infinigen: Everything From Scratch

Infinigen synthesizes natural-world scenes procedurally — terrain, plants, creatures, weather, water — optimized specifically for computer vision research. Built on Blender, it automatically generates high-quality ground truth including depth, surface normals, occlusion boundaries, semantic and instance segmentation, and bounding boxes.

A practical sense of the workflow — install, then point it at a scene count:

pip install infinigen
python -m infinigen.main --singlegpu --output_folder out/train --num_scenes 10

Three properties matter for training:

  • No assets are shared between Infinigen's training and evaluation data — a genuinely meaningful claim for generalization testing, since a model can't succeed merely by recognizing a specific 3D asset it saw during training, because every instance is generated fresh from the underlying rules and parameters.
  • Because the system exposes the procedural rules and parameters underlying each 3D scene, researchers get fine-grained control — a deeper version of the domain randomization idea from the previous article: randomizing not just the lighting and placement of fixed objects but the shape parameters of the objects themselves.
  • Demonstrated applications include producing 30K image pairs for training rectified stereo matching, alongside object detection, semantic segmentation, optical flow, and 3D reconstruction — and notably, a model trained on Infinigen data was reported to outperform one trained on existing synthetic datasets for at least some of these tasks.

Infinigen Indoors and the Constraint Solver

The original Infinigen covered only natural scenes. Infinigen Indoors extended it to the genuinely harder domain of indoor environments, where plausibility depends on layout logic — a bed shouldn't be inside a wall, a dining table needs chairs around it, a lamp belongs near something.

The system includes a constraint language and a constraint solver that searches for valid object arrangements satisfying stated requirements, letting users express generation objectives declaratively rather than hand-writing placement logic for every scene type. The conceptual parallel to the SDV tutorial's inequality constraint is exact — "ship date must not precede order date" there, "chairs must face the table" here: encode domain rules as explicit constraints the generator must satisfy, rather than hoping a statistical model infers them implicitly.

Infinigen Indoors generates every asset procedurally with no professional designers required, in contrast to ProcTHOR's approach below — a genuine difference in how unlimited the resulting asset diversity actually is.

ProcTHOR: Procedural Houses for Embodied AI

ProcTHOR (AI2, NeurIPS 2022) takes a different, heuristic-driven approach to a related goal — generating large numbers of interactive house layouts specifically for training embodied agents inside the AI2-THOR simulator.

It generates house floor plans and populates them using heuristics, drawing on an asset library built by professional designers — meaning its per-object diversity is bounded by that library's roughly 1.6K assets, even though the number of distinct house layouts it can compose from them is effectively unlimited.

This matters directly for the robotics material earlier in this series — the mobile robot path planning article's indoor navigation training and the curriculum learning article's environment-complexity progression both lean on procedurally generated houses to obtain thousands of distinct training environments without hand-building each one.

A genuine contrast with Infinigen worth stating plainly: ProcTHOR trades unlimited asset diversity for faster integration with an existing agent-training simulator; Infinigen trades that convenience for fully-procedural assets and photorealistic rendering suited to perception tasks.

Does It Actually Work? Evidence, With the Honest Nuance

The "combine real and synthetic, don't replace" principle from the earlier synthetic data work is a concrete instance here.

The Infinigen Indoors paper tested the data on shadow removal and occlusion boundary detection — two tasks genuinely lacking abundant existing training data, where synthetic data's value is easiest to demonstrate since there's less real data to compete against.

Combining real data with roughly 2K Infinigen Indoors synthetic images improved generalization performance relative to training on real data alone for these data-limited tasks.

The honest nuance, stated in the paper's own results: training on synthetic data alone led to slightly worse performance than real data on at least one comparison — the benefit came specifically from the real-plus-synthetic combination (their "R+S" condition), not from replacing real data entirely. The same pattern the synthetic data fundamentals article describes for GAN and diffusion-generated data applies here: synthetic scenes are most reliably valuable as a supplement anchoring or extending real data, not a wholesale substitute for it.

The Honest Limitation Current Research Names Directly

A September 2026 paper states a genuine, specific weakness plainly: procedural approaches like ProcTHOR and Infinigen address scalability and physical soundness, but tend to produce sterile, homogeneous layouts skewed toward simple rectangular rooms.

This is a real, distinct failure mode from anything the GAN and diffusion articles covered — procedural scenes are physically valid and perfectly labeled, but the distribution of layouts they produce can be narrower and tidier than genuinely lived-in real-world spaces, with their clutter, irregular geometry, and idiosyncratic arrangements.

A model trained on overly clean procedural scenes can still overfit to that cleanliness — adding realistic clutter and visual noise is the counter-measure, and it applies here as much as to hand-authored simulation.

Recent work specifically targets this gap by combining procedural methods with learned ones: VLM-based and LLM-guided 3D scene generation (graph-validated optimization for VLM-driven indoor scene generation is a named 2026 ECCV direction) uses a vision-language model's broad world knowledge to propose more realistic, varied layouts, while a procedural or optimization backend keeps them physically valid.

Where This Fits Against Generative Approaches

The same "combine rather than choose" pattern shows up here directly:

  • Procedural generation supplies geometry with guaranteed-correct labels — exact bounding boxes, depth, segmentation, from the scene graph itself.
  • Generative models supply appearance realism — diffusion-generated textures and backgrounds can be layered onto procedurally generated geometry, exactly the pipeline pattern the earlier decision framework described.
  • Neither family alone covers both needs — a purely generative image has no inherent ground-truth geometry, while a purely procedural scene may look too sterile; combining both is where current research is genuinely heading.

A Practical Decision Framework

  1. Do you need perfect geometric ground truth — depth, normals, exact segmentation — rather than just plausible-looking images? If yes, procedural generation (or hand-authored simulation) beats purely generative approaches structurally, since labels come from known scene geometry rather than inference.
  2. Is asset diversity itself your bottleneck? If a finite asset library is capping your diversity, fully-procedural assets (Infinigen) address that; if you mainly need many environment layouts for agent training, ProcTHOR's integration with an existing simulator may serve you better.
  3. Is your target domain natural scenes, or indoor environments? Infinigen's original system covers natural-world generation; Infinigen Indoors extends to indoor spaces with its constraint solver — pick the variant matching your actual deployment domain.
  4. Will real data be available to combine with the synthetic scenes? The R+S evidence supports combining both, so plan for a real-data component rather than assuming synthetic-only training will match it.
  5. Is realistic visual messiness important for your task? Plan for deliberate clutter, distractor objects, and possibly generative-model texture realism layered on top, rather than trusting clean procedural output alone.

Common Mistakes People Make

  • Treating procedural generation as interchangeable with hand-authored simulation. Procedural generation creates the assets and layouts themselves from rules, rather than randomizing placement of pre-made ones, and that difference is what removes the finite-asset-library ceiling.
  • Training on synthetic scenes alone and expecting real-world parity. The Infinigen Indoors results say the gains came from combining real and synthetic data, and synthetic-only performance slightly lagged on at least one comparison.
  • Ignoring layout homogeneity in procedurally generated scenes. Procedural outputs skew toward tidy, rectangular, sterile spaces, which can limit transfer to genuinely cluttered real environments.
  • Skipping explicit constraint design for indoor scenes. Without encoding layout logic like Infinigen Indoors' solver does, generated indoor scenes risk physically implausible object arrangements that teach a model unrealistic patterns.
  • Picking a tool by reputation rather than by whether it generates the domain you actually need. The original Infinigen is built for natural scenes, not interiors, and ProcTHOR is built for embodied-agent training rather than perception-quality rendering.
CoverBookDescriptionGet it
Cover of “Procedural Content Generation for Games” Procedural Content Generation for Gamesby Shaker, Togelius, and Nelson the design-side foundation: grammars, constraint-based generation, and why rule spaces beat asset libraries for variety. View on Amazon
Cover of “Mathematics for 3D Game Programming and Computer Graphics” Mathematics for 3D Game Programming and Computer Graphicsby Eric Lengyel the geometry and rendering math that procedural generators like Infinigen are built on top of inside Blender. View on Amazon
Cover of “Blender 3D” Blender 3Dby Example practical Blender authoring, since Infinigen sits on Blender and understanding its scene graph makes the ground-truth outputs (depth, normals, segmentation) far easier to interpret. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What's the difference between procedural generation and domain randomization?

Domain randomization varies lighting, textures, camera angles, and object placement within a scene built from pre-made 3D assets — the underlying sofa, room, and tree still came from an asset library or a team of 3D artists. Procedural generation writes rules that create the geometry, materials, and layout themselves, so diversity comes from the rule space rather than from a finite asset library. It's a deeper version of the same idea: randomizing the shape parameters of objects, not just their placement.

How do Infinigen and ProcTHOR differ?

Infinigen generates every asset from scratch using mathematical rules — no pre-made assets at all — which makes asset diversity effectively unlimited and gives it photoreal rendering suited to perception tasks, with no assets shared between its training and evaluation data. ProcTHOR takes a heuristic-driven approach on top of an asset library built by professional designers (roughly 1.6K assets), trading per-object diversity for fast integration with the AI2-THOR agent-training simulator and effectively unlimited house layout composition.

Does training on procedural synthetic data alone work?

The evidence says combine rather than substitute. The Infinigen Indoors paper found that adding roughly 2K synthetic images to real data improved generalization on data-limited tasks like shadow removal and occlusion boundary detection, while training on synthetic data alone was slightly worse than real data on at least one comparison. Synthetic scenes are most reliably valuable as a supplement anchoring or extending real data, not a wholesale replacement for it.

What is the main limitation of procedural scenes?

Layout homogeneity. A September 2026 paper states plainly that procedural approaches like ProcTHOR and Infinigen address scalability and physical soundness but tend to produce sterile, homogeneous layouts skewed toward simple rectangular rooms — physically valid and perfectly labeled, but tidier than lived-in real spaces with their clutter and irregular geometry. A model trained on overly clean procedural scenes can overfit to that cleanliness, so deliberate clutter and generative texture realism are the counter-measures.

Wrapping Up

Procedural scene generation attacks the asset bottleneck underlying the previous article's simulation-based approach — instead of sourcing or modeling 3D content, Infinigen generates everything from mathematical rules, producing unlimited, fully-labeled variety with no shared assets between training and test, while ProcTHOR trades that unlimited asset diversity for tighter integration with an embodied-agent simulator. Infinigen Indoors' constraint solver conceptually mirrors the SDV constraints covered earlier in this series, applying the same "encode domain rules explicitly" principle to 3D layout rather than tabular columns.

Remember that the evidence supports combining procedural synthetic data with real data rather than substituting it, and that current research is explicit about procedural scenes' tendency toward sterile, homogeneous layouts — a real limitation motivating the field's move toward hybrid systems pairing procedural validity with learned, VLM-driven realism.

Now go look back at the terrain-difficulty curriculum from the quadruped locomotion article and notice that it's, in effect, a small procedural generator with a difficulty parameter — then consider how a similar rule-based generator, with a clutter or complexity knob, could produce graded training scenes for whatever vision task you're working on. That reframing, more than any specific tool in this article, is the practical takeaway worth carrying forward.

What are You Looking For?

esc