# LEGO-Anything: Coding Agents for 3D Scene Reconstruction

> LEGO-Anything turns single images into editable 3D scene programs via coding agents, but current models achieve only about half the fidelity of specialist vision systems.

- **Source:** [arXiv](https://arxiv.org/abs/2609.36380)
- **Published:** 2026-10-01
- **Permalink:** https://picx.dev/p/sxMqJv
- **Whiteboard:** https://picx.dev/p/sxMqJv/image

## Summary

## Summary (Overview)

- **Framework Introduction**: The paper presents LEGO-Anything, an Image-to-Code (Image2Code) framework where coding agents reconstruct 3D scenes from single images by iteratively writing, executing, and refining Blender code, producing editable and executable scene programs.
- **Benchmark Creation**: The authors introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse scenes (indoor/outdoor, 8 environments, 17 themes), designed for precise automatic evaluation of end-to-end scene reconstruction.
- **Key Findings**: GPT-6-astra achieves the strongest results (53.4% indoor, 39.6% outdoor scores), but significant gaps remain between artifact validity and geometric/visual fidelity. Three recurring failure modes are identified: weak scene initialization, regressive edits, and unreliable self-assessment.
- **Plugin Solution**: LEGO-Plugin, a training-free harness plugin with three modules (Enhanced Initialization, Grounded Refinement, and a third targeting identified failures), improves every model without training, with the largest gains for weaker agents.
- **Downstream Evaluation**: Reconstructed scenes support detection, segmentation, and depth estimation as deterministic readouts, achieving about half of specialist box AP but falling well short of specialized vision models.

## Introduction and Theoretical Foundation

The paper addresses a fundamental question in 3D reconstruction: what representation is most useful? The authors argue that an executable scene program—which makes objects, geometry, layout, and camera explicit—is superior to meshes, point maps, or object sets because it can be run, inspected, edited, and queried like code.

The theoretical foundation builds on three research directions:
1. **Program-based approaches**: Recent work (Hu et al., 2024; Yin et al., 2026) treats scene reconstruction as program construction rather than direct 3D output prediction.
2. **Coding agents**: General-purpose agents combining language models with tools for code editing, execution, and iterative verification (Anthropic, 2026; Ning et al., 2026).
3. **Executable scene programs**: The representation makes reconstruction testable and diagnosable, unlike fixed 3D outputs.

The paper poses three central questions:
1. How faithfully do coding agents reconstruct scenes from a single image?
2. What limits their iterative construction process?
3. Are the resulting scene programs precise enough for downstream tasks?

## Methodology

### LEGO-Anything Framework
The framework formalizes reconstruction as:
- **Input**: Single RGB image $I$
- **Process**: A coding agent $\pi$ interacts with a 3D editing environment $E$ to produce a scene program $P$
- **Output**: Executed scene $S$ with construction trajectory $\tau = \{ (P_t, S_t, o_t) \}_{t=1}^{T}$

Four blocks define the process:
1. **Agent Workspace**: Manages artifacts
2. **Action Space**: Exposes agent operations
3. **Blender Code Structure**: Organizes editable scene state
4. **Scene Construction Workflow**: Iterates from reference image to final scene

### LEGO-Bench Benchmark
Built in LychSim using Fab assets, the benchmark features:
- 208 RGB inputs from 104 scenes
- 8 environments, 17 themes, 443 registered assets
- Nested difficulty levels: $\text{Easy} \subset \text{Medium} \subset \text{Hard}$
- Fixed architecture, materials, lighting, and camera per theme while visible content increases

### Evaluation Metrics
Three complementary axes:
- **Validity** $V_i$: Checks scene.blend opens, scene.glb is well-formed, final.png is parseable
- **Reconstruction** $R_i$: Visible-surface geometry compared against reference
- **Appearance** $A_i$: Rendered appearance fidelity

### LEGO-Plugin
Training-free harness plugin with three modules:
1. **Enhanced Initialization**: Recovers scene frame and camera using VGGT, derives visible-object layout cues
2. **Grounded Refinement**: Replaces free-form self-correction with tool-grounded measurement
3. **Third module**: Targets unreliable iteration (details in paper)

## Empirical Validation / Results

### Main Results (Table 3)
- GPT-6-astra achieves strongest overall: 53.4% indoor, 39.6% outdoor
- Coding agents deliver valid artifacts across both splits with stable Validity
- Fidelity drops outdoors for every model

### Key Findings

**Finding 2**: Greater scene complexity lowers Reconstruction and Appearance scores while leaving Validity intact. Outdoor scenes are harder than indoor at every complexity tier.

**Finding 4**: Reconstruction quality is non-monotonic over construction trajectories—weaker agents take longer to produce evaluable scenes, and subsequent edits can degrade fidelity.

**Finding 5**: Coding agents cannot reliably judge scene quality, especially geometry. Self-judging offers no advantage over cross-model judging.

**Finding 6**: LEGO-Plugin improves every model without training, with largest gains for weaker agents.

### LEGO-World Downstream Evaluation
Testing on 100 random images from COCO val2017, LVIS v1 val, and ETH3D:

| Task | Performance |
|------|------------|
| Detection | ~50% of specialist box AP |
| Segmentation | Non-trivial but below SOTA |
| Depth | Usable but imprecise |

**Finding 7**: Reconstructed scenes support detection, segmentation, and depth without task-specific training but fall well short of specialized models.

## Theoretical and Practical Implications

The paper's contributions have several significant implications:

1. **Representation Choice**: Executable scene programs provide a measurable, diagnosable target for reconstruction, unlike fixed 3D outputs.

2. **Agent Limitations**: The identification of three recurring failure modes (initialization, regressive edits, unreliable self-assessment) provides clear targets for improving coding agents.

3. **Evaluation Methodology**: LEGO-Bench's simulator-grounded design offers extensible, precise evaluation that scales with asset libraries—new environments, difficulty levels, and views can be added without changing the evaluation protocol.

4. **Plugin Approach**: LEGO-Plugin demonstrates that training-free interventions can meaningfully improve agent performance, suggesting a path forward that doesn't require model retraining.

5. **Downstream Usability**: The finding that frozen reconstructed scenes support multiple vision tasks positions executable scenes as a promising intermediate representation, though current precision limits practical deployment.

## Conclusion

The paper presents LEGO-Anything as an Image-to-Code framework where coding agents iteratively reconstruct 3D scenes as editable, executable programs. LEGO-Bench's three-axis evaluation (validity, reconstruction, appearance) reveals that current agents reliably deliver valid scene artifacts but remain limited in geometric fidelity. LEGO-Plugin successfully addresses the identified failure modes, improving all models without training.

The authors position executable scene programs as a promising, measurable, and diagnosable target for general-purpose coding agents. However, the substantial gaps between current performance and specialized models—particularly in geometric fidelity and downstream task support—highlight that program-constructed scenes from current coding agents are promising but not yet sufficiently precise for practical applications.

**Future directions** include:
- Improving geometric fidelity through better initialization and grounded refinement
- Extending LEGO-Bench with more environments and difficulty levels
- Exploring whether larger or differently-trained agents can close the remaining gap
- Investigating hybrid approaches combining program-based and learned reconstruction methods

---

_Markdown view of https://picx.dev/p/sxMqJv, served by PicX — AI-generated visual whiteboard summaries of research papers._
