Summary (Overview)

  • Framework Introduction: The paper presents LEGO-Anything, an Image-to-Code (Image2Code) framework where coding agents reconstruct 3D scenes from single images by iteratively writing, executing, and refining Blender code, producing editable and executable scene programs.
  • Benchmark Creation: The authors introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse scenes (indoor/outdoor, 8 environments, 17 themes), designed for precise automatic evaluation of end-to-end scene reconstruction.
  • Key Findings: GPT-6-astra achieves the strongest results (53.4% indoor, 39.6% outdoor scores), but significant gaps remain between artifact validity and geometric/visual fidelity. Three recurring failure modes are identified: weak scene initialization, regressive edits, and unreliable self-assessment.
  • Plugin Solution: LEGO-Plugin, a training-free harness plugin with three modules (Enhanced Initialization, Grounded Refinement, and a third targeting identified failures), improves every model without training, with the largest gains for weaker agents.
  • Downstream Evaluation: Reconstructed scenes support detection, segmentation, and depth estimation as deterministic readouts, achieving about half of specialist box AP but falling well short of specialized vision models.

Introduction and Theoretical Foundation

The paper addresses a fundamental question in 3D reconstruction: what representation is most useful? The authors argue that an executable scene program—which makes objects, geometry, layout, and camera explicit—is superior to meshes, point maps, or object sets because it can be run, inspected, edited, and queried like code.

The theoretical foundation builds on three research directions:

  1. Program-based approaches: Recent work (Hu et al., 2024; Yin et al., 2026) treats scene reconstruction as program construction rather than direct 3D output prediction.
  2. Coding agents: General-purpose agents combining language models with tools for code editing, execution, and iterative verification (Anthropic, 2026; Ning et al., 2026).
  3. Executable scene programs: The representation makes reconstruction testable and diagnosable, unlike fixed 3D outputs.

The paper poses three central questions:

  1. How faithfully do coding agents reconstruct scenes from a single image?
  2. What limits their iterative construction process?
  3. Are the resulting scene programs precise enough for downstream tasks?

Methodology

LEGO-Anything Framework

The framework formalizes reconstruction as:

  • Input: Single RGB image II
  • Process: A coding agent π\pi interacts with a 3D editing environment EE to produce a scene program PP
  • Output: Executed scene SS with construction trajectory τ={(Pt,St,ot)}t=1T\tau = \{ (P_t, S_t, o_t) \}_{t=1}^{T}

Four blocks define the process:

  1. Agent Workspace: Manages artifacts
  2. Action Space: Exposes agent operations
  3. Blender Code Structure: Organizes editable scene state
  4. Scene Construction Workflow: Iterates from reference image to final scene

LEGO-Bench Benchmark

Built in LychSim using Fab assets, the benchmark features:

  • 208 RGB inputs from 104 scenes
  • 8 environments, 17 themes, 443 registered assets
  • Nested difficulty levels: Easy⊂Medium⊂Hard\text{Easy} \subset \text{Medium} \subset \text{Hard}
  • Fixed architecture, materials, lighting, and camera per theme while visible content increases

Evaluation Metrics

Three complementary axes:

  • Validity ViV_i: Checks scene.blend opens, scene.glb is well-formed, final.png is parseable
  • Reconstruction RiR_i: Visible-surface geometry compared against reference
  • Appearance AiA_i: Rendered appearance fidelity

LEGO-Plugin

Training-free harness plugin with three modules:

  1. Enhanced Initialization: Recovers scene frame and camera using VGGT, derives visible-object layout cues
  2. Grounded Refinement: Replaces free-form self-correction with tool-grounded measurement
  3. Third module: Targets unreliable iteration (details in paper)

Empirical Validation / Results

Main Results (Table 3)

  • GPT-6-astra achieves strongest overall: 53.4% indoor, 39.6% outdoor
  • Coding agents deliver valid artifacts across both splits with stable Validity
  • Fidelity drops outdoors for every model

Key Findings

Finding 2: Greater scene complexity lowers Reconstruction and Appearance scores while leaving Validity intact. Outdoor scenes are harder than indoor at every complexity tier.

Finding 4: Reconstruction quality is non-monotonic over construction trajectories—weaker agents take longer to produce evaluable scenes, and subsequent edits can degrade fidelity.

Finding 5: Coding agents cannot reliably judge scene quality, especially geometry. Self-judging offers no advantage over cross-model judging.

Finding 6: LEGO-Plugin improves every model without training, with largest gains for weaker agents.

LEGO-World Downstream Evaluation

Testing on 100 random images from COCO val2017, LVIS v1 val, and ETH3D:

TaskPerformance
Detection~50% of specialist box AP
SegmentationNon-trivial but below SOTA
DepthUsable but imprecise

Finding 7: Reconstructed scenes support detection, segmentation, and depth without task-specific training but fall well short of specialized models.

Theoretical and Practical Implications

The paper's contributions have several significant implications:

  1. Representation Choice: Executable scene programs provide a measurable, diagnosable target for reconstruction, unlike fixed 3D outputs.

  2. Agent Limitations: The identification of three recurring failure modes (initialization, regressive edits, unreliable self-assessment) provides clear targets for improving coding agents.

  3. Evaluation Methodology: LEGO-Bench's simulator-grounded design offers extensible, precise evaluation that scales with asset libraries—new environments, difficulty levels, and views can be added without changing the evaluation protocol.

  4. Plugin Approach: LEGO-Plugin demonstrates that training-free interventions can meaningfully improve agent performance, suggesting a path forward that doesn't require model retraining.

  5. Downstream Usability: The finding that frozen reconstructed scenes support multiple vision tasks positions executable scenes as a promising intermediate representation, though current precision limits practical deployment.

Conclusion

The paper presents LEGO-Anything as an Image-to-Code framework where coding agents iteratively reconstruct 3D scenes as editable, executable programs. LEGO-Bench's three-axis evaluation (validity, reconstruction, appearance) reveals that current agents reliably deliver valid scene artifacts but remain limited in geometric fidelity. LEGO-Plugin successfully addresses the identified failure modes, improving all models without training.

The authors position executable scene programs as a promising, measurable, and diagnosable target for general-purpose coding agents. However, the substantial gaps between current performance and specialized models—particularly in geometric fidelity and downstream task support—highlight that program-constructed scenes from current coding agents are promising but not yet sufficiently precise for practical applications.

Future directions include:

  • Improving geometric fidelity through better initialization and grounded refinement
  • Extending LEGO-Bench with more environments and difficulty levels
  • Exploring whether larger or differently-trained agents can close the remaining gap
  • Investigating hybrid approaches combining program-based and learned reconstruction methods

Related papers