Summary (Overview)
- AgentGarten is a framework for building "code worlds" — interactive environments that combine programmable simulators/game engines with a shared, real-time neural renderer, enabling agents to learn through visual-only exploration and interaction.
- The framework separates world state and rules (maintained by an engine, exported as structured conditions like depth or surface normals) from appearance (synthesized by a neural renderer from pretrained video models).
- Adversarial Forcing, the key technical contribution, enables distillation of a bidirectional video model into a real-time, few-step block-causal renderer via: (1) distribution matching on self-rollouts, (2) exact bitwise replay for gradient propagation through history, and (3) a real-data adversarial objective with exact R1/R2 regularization that avoids double backward passes.
- Agents perceive the world strictly through rendered first-person frames, act via short Python programs, and improve across rounds by writing playbooks (skill files) inherited by subsequent agents. In hide-and-seek, shelters emerged by round 4 and ramp use by round 10 — versus ~25M and ~100M episodes for tabula-rasa reinforcement learning.
- The renderer achieves over 35 frames per second at 480×832 resolution on a single H100 GPU, with a throughput of 438.8 ms per served block of 16 frames.
Introduction and Theoretical Foundation
Background and Motivation
The central challenge in agent training is that what agents can learn is bounded by the environments they practice in. Two requirements are essential:
- Faithfulness: Environments must have consistent state, rules, and dynamics — the consequences of past actions must persist.
- Realism: Observations must follow real-world visual distributions.
The Fundamental Trade-off
| Approach | Strengths | Weaknesses |
|---|---|---|
| Simulators / game engines | Explicit state, programmable rules, inspectable dynamics | Costly 3D asset creation, complex rendering pipelines |
| Video world models | Rich visual observations from data | Implicit state, non-inspectable rules, no persistence guarantees |
Key Insight
AgentGarten resolves this trade-off by separating world logic from appearance. A scene program runs in an engine that maintains persistent state; a shared neural renderer synthesizes visual observations from structured conditions.
Formal Formulation
The code world is defined by a state transition function and a condition-rendering function (Eq. 1):
where is scene state, is camera pose, is the agent's action, and is a structured condition (depth or surface normals). The neural renderer generates observations (Eq. 2):
Here is an appearance reference image, is a text description, and is visual history.
Critical property: "the transition takes no rendered observation as input: the renderer influences the state only through the actions that agents select, so rendering errors cannot accumulate in the state."
Methodology
1. Architecture Overview
- Backbone: Initialized from Cosmos 3-Nano [21], with frozen understanding (UND) and generation (GEN) towers. A frozen Wan video autoencoder maps RGB and geometry videos to latents.
- Geometry conditioning: Depth or surface normals are encoded via a frozen video encoder with spatial pooling (Eq. 3):
- Joint attention: Geometry and RGB tokens are concatenated along the sequence dimension (Eq. 4):
2. Teacher-Forcing Adaptation (Stage 2)
The model is adapted to blockwise autoregressive generation. Training uses parallel clean/noisy block copies with an attention mask (Eq. 5):
The loss (Eq. 6) trains all blocks in parallel with independently sampled flow times:
3. Adversarial Forcing (Stage 3)
Three components distinguish Adversarial Forcing:
(a) Distribution Matching (DMD)
Student rollouts are matched to the bidirectional teacher via flow-matching distillation (Eq. 7):
(b) Exact Replay
The key innovation for history gradients: a no-gradient rollout records each block's input and clean output ; a differentiable replay recomputes predictions with bitwise-identical execution (Eq. 8):
Table 1: Replay accuracy and cost (H100, 36 layers, 480×832, 61 latent frames, BF16):
| Measure | SGF-style (FlexAttention) | Ours (block-by-block SDPA) | Change |
|---|---|---|---|
| Replay error (relative ) | 3.99% | 0 (bitwise) | — |
| Forward time | 2.95 s | 2.79 s | −5.4% |
| Peak allocated memory | 50.83 GiB | 50.92 GiB | +0.2% |
(c) Adversarial Training with Exact R1/R2
A discriminator head on the frozen teacher backbone. The relativistic objective (Eq. 9):
The R1/R2 penalties (Eq. 10) are computed exactly without double backward through fused attention kernels, using the fact that the backbone is frozen. With , , (Eq. 11):
The parameter gradient (Eq. 12) is obtained via a Jacobian–vector product surrogate (Eq. 13):
4. Streaming Inference
- Bounded KV cache with sink prefix + sliding window
- Hand-written Triton kernels for fused elementwise ops (RMSNorm, RoPE, gated activations)
- CUDA graph capture eliminates host launch overhead
- Distilled tiny VAE decoder (<10 ms)
Table 2: Inference cost per served block (16 frames, one H100, BF16)
| Component | Time |
|---|---|
| Condition encoding | 10.6 ms |
| Transformer (4 denoising steps + cache) | 414.4 ms |
| Tiny decoder + host transfer | 8.8 ms |
| Total wall time | 438.8 ms |
| Throughput | 36.5 frames/s |
Empirical Validation / Results
4.1 Self Forcing vs. Adversarial Forcing
In 30-second rollouts on identical condition trajectories, Self Forcing develops repetitive surface patterns and loses scene detail over time. Adversarial Forcing retains natural textures and fine detail throughout (Figure 8).
4.2 Hide-and-Seek: Emergent Tool Use
Agents perceive only through neural-rendered camera frames, act via short Python programs, and maintain per-role playbooks. Results:
| Strategy | Self-play RL [13] episodes | AgentGarten rounds |
|---|---|---|
| Build shelters | ≈25 million | 4 |
| Use ramps to cross walls | ≈100 million | 10 |
Key emergent behaviors observed:
- A hider builds cover by moving a barrier panel to block an entrance
- A seeker transports a ramp to an inner wall and climbs over it
- A seeker repositions a ramp closer to the wall and retries after an initial failed leap
"These emergent strategies mirror the hallmark milestones of the 2019 study... The primary challenge is not discovering concepts from nothing, but grounding abstract knowledge into closed-loop sensorimotor action."
4.3 Additional Worlds (4 rounds each)
| World | Measure | Round 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|
| Companion dog | engagement score | 13 | 14 | 19 | 19 |
| One-lane bridge | seconds until both arrive | 71 | 68 | 45 | 41 |
| Herding | score / 100 | 60 | 90.1 | 87.6 | 88.3 |
| Quarry loader | score / 100 | 30 | 0 | 30 | 90.9 |
Theoretical and Practical Implications
Theoretical Significance
-
Exact replay vs. numerical approximation: The bitwise-identical replay (Table 1) resolves a subtle numerical divergence issue in prior Self Gradient Forcing approaches, enabling true gradient flow through history encoding without memory explosion.
-
Exact R1/R2 with frozen backbones: The paper derives a method to compute exact gradient penalties for discriminators with frozen fused-attention backbones, sidestepping double-backward limitations of FlashAttention — a contribution beyond this specific architecture.
-
The code worlds paradigm: Environments become executable, inspectable, and revisable programs rather than fixed assets. A recorded rollout can be re-rendered from a different camera or style without altering the events.
Practical Implications
- Learning efficiency: Foundation-model agents bypass tabula-rasa exploration, using web-scale priors that only need grounding — reducing effective learning time from tens of millions of episodes to 4–10 practice rounds.
- Scalability: New worlds are written as code and rendered through the same interface, so environments can grow in number and difficulty alongside agents.
- Real-time interaction: The renderer sustains >35 FPS on one H100, making interactive deployment feasible (browser games, multi-agent racing, etc.).
Conclusion
AgentGarten demonstrates a viable path toward agents that continue evolving through interactive experience. The key thesis: what an agent can learn is bounded by its environment, and code worlds with a shared neural renderer provide a scalable way to create environments that are both faithful (explicit state and rules) and realistic (neural-rendered observations).
Future Directions
-
Richer condition interfaces: Geometry alone cannot represent rotationally symmetric objects spinning, colors, or object identity. Structured text describing object attributes or high-dimensional latent features could extend the interface.
-
Automatic world validation: "The harder question is whether it deserves an agent's time: whether the task can be solved from what the agent sees, whether a careless strategy fails, and whether there is something to learn."
-
Scaling experience across worlds: Playbooks currently belong to one world; cross-world knowledge consolidation (merging, dropping, transferring lessons) remains open.
-
Agentic world authorship: Since worlds are programs, coding agents could write and revise new worlds, closing the loop where "where agents fail tells us which worlds to write next, and each new world tests what they wrote down before."
Related papers
- Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
A fixed LLM judge produces version-dependent errors, invalidating agent comparisons and transported calibration, so release decisions require paired audits, not judge-only scores.
- LLM sequential decision making under uncertainty in biochemical domains
LLMs overreact to new data and fail to explore in Bayesian optimization due to context-stickiness, a competence gap, not an intention gap, fixable by removing in-context history.
- From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining
Search scaling in autonomous research is capability-dependent: deeper search narrows cross-model gaps, but early trajectory state and parallel exploration critically shape final performance.