# AgentGarten: Code Worlds for Evolving Agents

> AgentGarten separates world state from neural-rendered appearance, enabling agents to learn emergent tool use in 4-10 rounds versus millions of reinforcement learning episodes.

- **Source:** [arXiv](https://arxiv.org/abs/2610.12374)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/68E2YO
- **Whiteboard:** https://picx.dev/p/68E2YO/image

## Summary

## Summary (Overview)

- **AgentGarten** is a framework for building "code worlds" — interactive environments that combine programmable simulators/game engines with a shared, real-time neural renderer, enabling agents to learn through visual-only exploration and interaction.
- The framework separates **world state and rules** (maintained by an engine, exported as structured conditions like depth or surface normals) from **appearance** (synthesized by a neural renderer from pretrained video models).
- **Adversarial Forcing**, the key technical contribution, enables distillation of a bidirectional video model into a real-time, few-step block-causal renderer via: (1) distribution matching on self-rollouts, (2) *exact bitwise replay* for gradient propagation through history, and (3) a real-data adversarial objective with exact R1/R2 regularization that avoids double backward passes.
- Agents perceive the world strictly through rendered first-person frames, act via short Python programs, and improve across rounds by writing **playbooks** (skill files) inherited by subsequent agents. In hide-and-seek, shelters emerged by round 4 and ramp use by round 10 — versus ~25M and ~100M episodes for tabula-rasa reinforcement learning.
- The renderer achieves **over 35 frames per second** at 480×832 resolution on a single H100 GPU, with a throughput of 438.8 ms per served block of 16 frames.

---

## Introduction and Theoretical Foundation

### Background and Motivation

The central challenge in agent training is that **what agents can learn is bounded by the environments they practice in**. Two requirements are essential:

1. **Faithfulness**: Environments must have consistent state, rules, and dynamics — the consequences of past actions must persist.
2. **Realism**: Observations must follow real-world visual distributions.

### The Fundamental Trade-off

| Approach | Strengths | Weaknesses |
|----------|-----------|------------|
| **Simulators / game engines** | Explicit state, programmable rules, inspectable dynamics | Costly 3D asset creation, complex rendering pipelines |
| **Video world models** | Rich visual observations from data | Implicit state, non-inspectable rules, no persistence guarantees |

### Key Insight

AgentGarten resolves this trade-off by **separating world logic from appearance**. A scene program $p$ runs in an engine that maintains persistent state; a shared neural renderer $R_\theta$ synthesizes visual observations from structured conditions.

### Formal Formulation

The code world is defined by a state transition function and a condition-rendering function (Eq. 1):

$$
(s_{t+1}, \pi_{t+1}) = f_p(s_t, \pi_t, a_t), \qquad c_t = h_p(s_t, \pi_t),\tag{1}
$$

where $s_t$ is scene state, $\pi_t$ is camera pose, $a_t$ is the agent's action, and $c_t$ is a structured condition (depth or surface normals). The neural renderer generates observations (Eq. 2):

$$
x_t = R_\theta(x_{<t}, c_{\leq t}, x_0, y), \qquad t \geq 1.\tag{2}
$$

Here $x_0$ is an appearance reference image, $y$ is a text description, and $x_{<t}$ is visual history.

> **Critical property**: "the transition $f_p$ takes no rendered observation as input: the renderer influences the state only through the actions that agents select, so rendering errors cannot accumulate in the state."

---

## Methodology

### 1. Architecture Overview

- **Backbone**: Initialized from Cosmos 3-Nano [21], with frozen understanding (UND) and generation (GEN) towers. A frozen Wan video autoencoder maps RGB and geometry videos to latents.
- **Geometry conditioning**: Depth or surface normals are encoded via a frozen video encoder with spatial pooling (Eq. 3):

$$G_t = W_{\text{in}} \mathcal{P}([E(S(c))]_t) + e_{\text{geo}},\tag{3}$$

- **Joint attention**: Geometry and RGB tokens are concatenated along the sequence dimension (Eq. 4):

$$[G_0, G_1, \dots][R_0, R_1, \dots],\tag{4}$$

### 2. Teacher-Forcing Adaptation (Stage 2)

The model is adapted to blockwise autoregressive generation. Training uses parallel clean/noisy block copies with an attention mask (Eq. 5):

$$P_j \longrightarrow A_0 \cup P_{\leq j}, \quad Q_j \longrightarrow A_0 \cup P_{<j} \cup Q_j,\tag{5}$$

The loss (Eq. 6) trains all blocks in parallel with independently sampled flow times:

$$z_{j,t_j} = (1 - t_j) z_j + t_j \epsilon_j, \qquad \epsilon_j \sim \mathcal{N}(0, I),$$

$$\mathcal{L}_{\text{TF}} = \mathbf{E}\left[ \frac{1}{M} \sum_{j=1}^{M} \left\| v_\theta(z_{j,t_j}, t_j \mid z_{<j}, \mathcal{C}) - (\epsilon_j - z_j) \right\|_2^2 \right],\tag{6}$$

### 3. Adversarial Forcing (Stage 3)

Three components distinguish Adversarial Forcing:

#### (a) Distribution Matching (DMD)
Student rollouts are matched to the bidirectional teacher via flow-matching distillation (Eq. 7):

$$g = \frac{f_\psi(y_\tau, \tau \mid \mathcal{C}) - f_{\mathrm{T}}(y_\tau, \tau \mid \mathcal{C})}{a}, \qquad \mathcal{L}_{\mathrm{DMD}} = \mathbf{E}\left[ \frac{1}{2N} \| \tilde{z}_\theta - \mathrm{sg}(\tilde{z}_\theta - g) \|_2^2 \right]\tag{7}$$

#### (b) Exact Replay
The key innovation for history gradients: a no-gradient rollout records each block's input $U_j$ and clean output $Z_j$; a differentiable replay recomputes predictions with bitwise-identical execution (Eq. 8):

$$\tilde{z}_{\theta,j} = \mathrm{sg}(U_j) - t^* v_\theta\big(\mathrm{sg}(U_j), t^* \mid \mathrm{sg}(Z_{<j}), \mathcal{C}; \mathcal{M}_{\mathrm{TF}}\big).\tag{8}$$

**Table 1: Replay accuracy and cost** (H100, 36 layers, 480×832, 61 latent frames, BF16):

| Measure | SGF-style (FlexAttention) | Ours (block-by-block SDPA) | Change |
|---------|---------------------------|----------------------------|--------|
| Replay error (relative $L_2$) | 3.99% | **0 (bitwise)** | — |
| Forward time | 2.95 s | **2.79 s** | −5.4% |
| Peak allocated memory | 50.83 GiB | 50.92 GiB | +0.2% |

#### (c) Adversarial Training with Exact R1/R2

A discriminator head $h_\phi$ on the frozen teacher backbone. The relativistic objective (Eq. 9):

$$\mathcal{L}_{\mathrm{G}} = \mathbf{E}[\mathrm{softplus}(\mathrm{sg}(r) - f)], \qquad \mathcal{L}_{\mathrm{rel}} = \mathbf{E}[\mathrm{softplus}(f - r)],\tag{9}$$

The R1/R2 penalties (Eq. 10) are computed **exactly** without double backward through fused attention kernels, using the fact that the backbone is frozen. With $z = B(x)$, $A = J_B(x)$, $u = \nabla_z h_\phi(z)$ (Eq. 11):

$$g = A^\top u, \qquad R = g^\top g, \qquad v = A g,\tag{11}$$

The parameter gradient (Eq. 12) is obtained via a Jacobian–vector product surrogate (Eq. 13):

$$\nabla_\phi R = 2\left(\frac{\partial u}{\partial \phi}\right)^\top v = \nabla_\phi\left[ 2 J_z h_\phi(z) \mathrm{sg}(v) \right],\tag{12}$$

$$\widetilde{R} = \mathrm{sg}(\|g\|_2^2) + s - \mathrm{sg}(s)\tag{13}$$

### 4. Streaming Inference

- Bounded KV cache with sink prefix + sliding window
- Hand-written Triton kernels for fused elementwise ops (RMSNorm, RoPE, gated activations)
- CUDA graph capture eliminates host launch overhead
- Distilled tiny VAE decoder (<10 ms)

**Table 2: Inference cost per served block (16 frames, one H100, BF16)**

| Component | Time |
|-----------|------|
| Condition encoding | 10.6 ms |
| Transformer (4 denoising steps + cache) | 414.4 ms |
| Tiny decoder + host transfer | 8.8 ms |
| **Total wall time** | **438.8 ms** |
| **Throughput** | **36.5 frames/s** |

---

## Empirical Validation / Results

### 4.1 Self Forcing vs. Adversarial Forcing

In 30-second rollouts on identical condition trajectories, Self Forcing develops repetitive surface patterns and loses scene detail over time. **Adversarial Forcing retains natural textures and fine detail throughout** (Figure 8).

### 4.2 Hide-and-Seek: Emergent Tool Use

Agents perceive only through neural-rendered camera frames, act via short Python programs, and maintain per-role playbooks. Results:

| Strategy | Self-play RL [13] episodes | AgentGarten rounds |
|----------|---------------------------|-------------------|
| Build shelters | ≈25 million | **4** |
| Use ramps to cross walls | ≈100 million | **10** |

Key emergent behaviors observed:
- A hider builds cover by moving a barrier panel to block an entrance
- A seeker transports a ramp to an inner wall and climbs over it
- A seeker repositions a ramp closer to the wall and retries after an initial failed leap

> "These emergent strategies mirror the hallmark milestones of the 2019 study... The primary challenge is not discovering concepts from nothing, but grounding abstract knowledge into closed-loop sensorimotor action."

### 4.3 Additional Worlds (4 rounds each)

| World | Measure | Round 1 | 2 | 3 | 4 |
|-------|---------|---------|---|---|---|
| Companion dog | engagement score | 13 | 14 | 19 | 19 |
| One-lane bridge | seconds until both arrive | 71 | 68 | 45 | **41** |
| Herding | score / 100 | 60 | 90.1 | 87.6 | 88.3 |
| Quarry loader | score / 100 | 30 | 0 | 30 | **90.9** |

---

## Theoretical and Practical Implications

### Theoretical Significance

1. **Exact replay vs. numerical approximation**: The bitwise-identical replay (Table 1) resolves a subtle numerical divergence issue in prior Self Gradient Forcing approaches, enabling true gradient flow through history encoding without memory explosion.

2. **Exact R1/R2 with frozen backbones**: The paper derives a method to compute exact gradient penalties for discriminators with frozen fused-attention backbones, sidestepping double-backward limitations of FlashAttention — a contribution beyond this specific architecture.

3. **The code worlds paradigm**: Environments become *executable, inspectable, and revisable programs* rather than fixed assets. A recorded rollout can be re-rendered from a different camera or style without altering the events.

### Practical Implications

- **Learning efficiency**: Foundation-model agents bypass tabula-rasa exploration, using web-scale priors that only need *grounding* — reducing effective learning time from tens of millions of episodes to 4–10 practice rounds.
- **Scalability**: New worlds are written as code and rendered through the same interface, so environments can grow in number and difficulty alongside agents.
- **Real-time interaction**: The renderer sustains >35 FPS on one H100, making interactive deployment feasible (browser games, multi-agent racing, etc.).

---

## Conclusion

AgentGarten demonstrates a viable path toward **agents that continue evolving through interactive experience**. The key thesis: what an agent can learn is bounded by its environment, and code worlds with a shared neural renderer provide a scalable way to create environments that are both faithful (explicit state and rules) and realistic (neural-rendered observations).

### Future Directions

1. **Richer condition interfaces**: Geometry alone cannot represent rotationally symmetric objects spinning, colors, or object identity. Structured text describing object attributes or high-dimensional latent features could extend the interface.

2. **Automatic world validation**: "The harder question is whether it deserves an agent's time: whether the task can be solved from what the agent sees, whether a careless strategy fails, and whether there is something to learn."

3. **Scaling experience across worlds**: Playbooks currently belong to one world; cross-world knowledge consolidation (merging, dropping, transferring lessons) remains open.

4. **Agentic world authorship**: Since worlds are programs, coding agents could write and revise new worlds, closing the loop where "where agents fail tells us which worlds to write next, and each new world tests what they wrote down before."

---

_Markdown view of https://picx.dev/p/68E2YO, served by PicX — AI-generated visual whiteboard summaries of research papers._
