Summary (Overview)

  • AgentGarten is a framework for building "code worlds" — interactive environments that combine programmable simulators/game engines with a shared, real-time neural renderer, enabling agents to learn through visual-only exploration and interaction.
  • The framework separates world state and rules (maintained by an engine, exported as structured conditions like depth or surface normals) from appearance (synthesized by a neural renderer from pretrained video models).
  • Adversarial Forcing, the key technical contribution, enables distillation of a bidirectional video model into a real-time, few-step block-causal renderer via: (1) distribution matching on self-rollouts, (2) exact bitwise replay for gradient propagation through history, and (3) a real-data adversarial objective with exact R1/R2 regularization that avoids double backward passes.
  • Agents perceive the world strictly through rendered first-person frames, act via short Python programs, and improve across rounds by writing playbooks (skill files) inherited by subsequent agents. In hide-and-seek, shelters emerged by round 4 and ramp use by round 10 — versus ~25M and ~100M episodes for tabula-rasa reinforcement learning.
  • The renderer achieves over 35 frames per second at 480×832 resolution on a single H100 GPU, with a throughput of 438.8 ms per served block of 16 frames.

Introduction and Theoretical Foundation

Background and Motivation

The central challenge in agent training is that what agents can learn is bounded by the environments they practice in. Two requirements are essential:

  1. Faithfulness: Environments must have consistent state, rules, and dynamics — the consequences of past actions must persist.
  2. Realism: Observations must follow real-world visual distributions.

The Fundamental Trade-off

ApproachStrengthsWeaknesses
Simulators / game enginesExplicit state, programmable rules, inspectable dynamicsCostly 3D asset creation, complex rendering pipelines
Video world modelsRich visual observations from dataImplicit state, non-inspectable rules, no persistence guarantees

Key Insight

AgentGarten resolves this trade-off by separating world logic from appearance. A scene program pp runs in an engine that maintains persistent state; a shared neural renderer RθR_\theta synthesizes visual observations from structured conditions.

Formal Formulation

The code world is defined by a state transition function and a condition-rendering function (Eq. 1):

(st+1,πt+1)=fp(st,πt,at),ct=hp(st,πt),(1)(s_{t+1}, \pi_{t+1}) = f_p(s_t, \pi_t, a_t), \qquad c_t = h_p(s_t, \pi_t),\tag{1}

where sts_t is scene state, πt\pi_t is camera pose, ata_t is the agent's action, and ctc_t is a structured condition (depth or surface normals). The neural renderer generates observations (Eq. 2):

xt=Rθ(x<t,c≤t,x0,y),t≥1.(2)x_t = R_\theta(x_{<t}, c_{\leq t}, x_0, y), \qquad t \geq 1.\tag{2}

Here x0x_0 is an appearance reference image, yy is a text description, and x<tx_{<t} is visual history.

Critical property: "the transition fpf_p takes no rendered observation as input: the renderer influences the state only through the actions that agents select, so rendering errors cannot accumulate in the state."


Methodology

1. Architecture Overview

  • Backbone: Initialized from Cosmos 3-Nano [21], with frozen understanding (UND) and generation (GEN) towers. A frozen Wan video autoencoder maps RGB and geometry videos to latents.
  • Geometry conditioning: Depth or surface normals are encoded via a frozen video encoder with spatial pooling (Eq. 3):
Gt=WinP([E(S(c))]t)+egeo,(3)G_t = W_{\text{in}} \mathcal{P}([E(S(c))]_t) + e_{\text{geo}},\tag{3}
  • Joint attention: Geometry and RGB tokens are concatenated along the sequence dimension (Eq. 4):
[G0,G1,… ][R0,R1,… ],(4)[G_0, G_1, \dots][R_0, R_1, \dots],\tag{4}

2. Teacher-Forcing Adaptation (Stage 2)

The model is adapted to blockwise autoregressive generation. Training uses parallel clean/noisy block copies with an attention mask (Eq. 5):

Pj⟶A0∪P≤j,Qj⟶A0∪P<j∪Qj,(5)P_j \longrightarrow A_0 \cup P_{\leq j}, \quad Q_j \longrightarrow A_0 \cup P_{<j} \cup Q_j,\tag{5}

The loss (Eq. 6) trains all blocks in parallel with independently sampled flow times:

zj,tj=(1−tj)zj+tjϵj,ϵj∼N(0,I),z_{j,t_j} = (1 - t_j) z_j + t_j \epsilon_j, \qquad \epsilon_j \sim \mathcal{N}(0, I), LTF=E[1M∑j=1M∥vθ(zj,tj,tj∣z<j,C)−(ϵj−zj)∥22],(6)\mathcal{L}_{\text{TF}} = \mathbf{E}\left[ \frac{1}{M} \sum_{j=1}^{M} \left\| v_\theta(z_{j,t_j}, t_j \mid z_{<j}, \mathcal{C}) - (\epsilon_j - z_j) \right\|_2^2 \right],\tag{6}

3. Adversarial Forcing (Stage 3)

Three components distinguish Adversarial Forcing:

(a) Distribution Matching (DMD)

Student rollouts are matched to the bidirectional teacher via flow-matching distillation (Eq. 7):

g=fψ(yτ,τ∣C)−fT(yτ,τ∣C)a,LDMD=E[12N∥z~θ−sg(z~θ−g)∥22](7)g = \frac{f_\psi(y_\tau, \tau \mid \mathcal{C}) - f_{\mathrm{T}}(y_\tau, \tau \mid \mathcal{C})}{a}, \qquad \mathcal{L}_{\mathrm{DMD}} = \mathbf{E}\left[ \frac{1}{2N} \| \tilde{z}_\theta - \mathrm{sg}(\tilde{z}_\theta - g) \|_2^2 \right]\tag{7}

(b) Exact Replay

The key innovation for history gradients: a no-gradient rollout records each block's input UjU_j and clean output ZjZ_j; a differentiable replay recomputes predictions with bitwise-identical execution (Eq. 8):

z~θ,j=sg(Uj)−t∗vθ(sg(Uj),t∗∣sg(Z<j),C;MTF).(8)\tilde{z}_{\theta,j} = \mathrm{sg}(U_j) - t^* v_\theta\big(\mathrm{sg}(U_j), t^* \mid \mathrm{sg}(Z_{<j}), \mathcal{C}; \mathcal{M}_{\mathrm{TF}}\big).\tag{8}

Table 1: Replay accuracy and cost (H100, 36 layers, 480×832, 61 latent frames, BF16):

MeasureSGF-style (FlexAttention)Ours (block-by-block SDPA)Change
Replay error (relative L2L_2)3.99%0 (bitwise)—
Forward time2.95 s2.79 s−5.4%
Peak allocated memory50.83 GiB50.92 GiB+0.2%

(c) Adversarial Training with Exact R1/R2

A discriminator head hϕh_\phi on the frozen teacher backbone. The relativistic objective (Eq. 9):

LG=E[softplus(sg(r)−f)],Lrel=E[softplus(f−r)],(9)\mathcal{L}_{\mathrm{G}} = \mathbf{E}[\mathrm{softplus}(\mathrm{sg}(r) - f)], \qquad \mathcal{L}_{\mathrm{rel}} = \mathbf{E}[\mathrm{softplus}(f - r)],\tag{9}

The R1/R2 penalties (Eq. 10) are computed exactly without double backward through fused attention kernels, using the fact that the backbone is frozen. With z=B(x)z = B(x), A=JB(x)A = J_B(x), u=∇zhϕ(z)u = \nabla_z h_\phi(z) (Eq. 11):

g=A⊤u,R=g⊤g,v=Ag,(11)g = A^\top u, \qquad R = g^\top g, \qquad v = A g,\tag{11}

The parameter gradient (Eq. 12) is obtained via a Jacobian–vector product surrogate (Eq. 13):

∇ϕR=2(∂u∂ϕ)⊤v=∇ϕ[2Jzhϕ(z)sg(v)],(12)\nabla_\phi R = 2\left(\frac{\partial u}{\partial \phi}\right)^\top v = \nabla_\phi\left[ 2 J_z h_\phi(z) \mathrm{sg}(v) \right],\tag{12} R~=sg(∥g∥22)+s−sg(s)(13)\widetilde{R} = \mathrm{sg}(\|g\|_2^2) + s - \mathrm{sg}(s)\tag{13}

4. Streaming Inference

  • Bounded KV cache with sink prefix + sliding window
  • Hand-written Triton kernels for fused elementwise ops (RMSNorm, RoPE, gated activations)
  • CUDA graph capture eliminates host launch overhead
  • Distilled tiny VAE decoder (<10 ms)

Table 2: Inference cost per served block (16 frames, one H100, BF16)

ComponentTime
Condition encoding10.6 ms
Transformer (4 denoising steps + cache)414.4 ms
Tiny decoder + host transfer8.8 ms
Total wall time438.8 ms
Throughput36.5 frames/s

Empirical Validation / Results

4.1 Self Forcing vs. Adversarial Forcing

In 30-second rollouts on identical condition trajectories, Self Forcing develops repetitive surface patterns and loses scene detail over time. Adversarial Forcing retains natural textures and fine detail throughout (Figure 8).

4.2 Hide-and-Seek: Emergent Tool Use

Agents perceive only through neural-rendered camera frames, act via short Python programs, and maintain per-role playbooks. Results:

StrategySelf-play RL [13] episodesAgentGarten rounds
Build shelters≈25 million4
Use ramps to cross walls≈100 million10

Key emergent behaviors observed:

  • A hider builds cover by moving a barrier panel to block an entrance
  • A seeker transports a ramp to an inner wall and climbs over it
  • A seeker repositions a ramp closer to the wall and retries after an initial failed leap

"These emergent strategies mirror the hallmark milestones of the 2019 study... The primary challenge is not discovering concepts from nothing, but grounding abstract knowledge into closed-loop sensorimotor action."

4.3 Additional Worlds (4 rounds each)

WorldMeasureRound 1234
Companion dogengagement score13141919
One-lane bridgeseconds until both arrive71684541
Herdingscore / 1006090.187.688.3
Quarry loaderscore / 1003003090.9

Theoretical and Practical Implications

Theoretical Significance

  1. Exact replay vs. numerical approximation: The bitwise-identical replay (Table 1) resolves a subtle numerical divergence issue in prior Self Gradient Forcing approaches, enabling true gradient flow through history encoding without memory explosion.

  2. Exact R1/R2 with frozen backbones: The paper derives a method to compute exact gradient penalties for discriminators with frozen fused-attention backbones, sidestepping double-backward limitations of FlashAttention — a contribution beyond this specific architecture.

  3. The code worlds paradigm: Environments become executable, inspectable, and revisable programs rather than fixed assets. A recorded rollout can be re-rendered from a different camera or style without altering the events.

Practical Implications

  • Learning efficiency: Foundation-model agents bypass tabula-rasa exploration, using web-scale priors that only need grounding — reducing effective learning time from tens of millions of episodes to 4–10 practice rounds.
  • Scalability: New worlds are written as code and rendered through the same interface, so environments can grow in number and difficulty alongside agents.
  • Real-time interaction: The renderer sustains >35 FPS on one H100, making interactive deployment feasible (browser games, multi-agent racing, etc.).

Conclusion

AgentGarten demonstrates a viable path toward agents that continue evolving through interactive experience. The key thesis: what an agent can learn is bounded by its environment, and code worlds with a shared neural renderer provide a scalable way to create environments that are both faithful (explicit state and rules) and realistic (neural-rendered observations).

Future Directions

  1. Richer condition interfaces: Geometry alone cannot represent rotationally symmetric objects spinning, colors, or object identity. Structured text describing object attributes or high-dimensional latent features could extend the interface.

  2. Automatic world validation: "The harder question is whether it deserves an agent's time: whether the task can be solved from what the agent sees, whether a careless strategy fails, and whether there is something to learn."

  3. Scaling experience across worlds: Playbooks currently belong to one world; cross-world knowledge consolidation (merging, dropping, transferring lessons) remains open.

  4. Agentic world authorship: Since worlds are programs, coding agents could write and revise new worlds, closing the loop where "where agents fail tells us which worlds to write next, and each new world tests what they wrote down before."

Related papers