Full text not available for this paper

Summary (Overview)

  • RSIAgent is a training-free multi-agent framework for recursive self-improvement in new digital environments, enabling agents to autonomously explore, verify, and consolidate environment-specific knowledge into reusable memory without updating model parameters.
  • The framework coordinates three agent roles: a curriculum agent (decides what to explore), an actor agent (executes actions and updates memory), and a verifier agent (grounds outcomes in environment feedback).
  • A broad-then-deep exploration strategy is introduced: Broad Recursive Self-exploration (BRS) builds diverse environment coverage in parallel, while Deep Recursive Self-exploration (DRS) focuses on hard cases, hidden constraints, and boundary conditions.
  • Experiments on OSWorld-v2 and Agent's Last Exam (ALE) show that RSIAgent enables open-source models (Kimi-K3, GLM-5.3) to outperform frontier closed-source models including GPT-6 Astra and Claude Opus 5.
  • Additional evaluation on GameCraft-Bench demonstrates generalization beyond computer-use tasks to interactive game development, consistently improving game quality across both weak and strong base generators.

Introduction and Theoretical Foundation

Background and Motivation

Large language models (LLMs) and vision-language models (VLMs) have advanced significantly in perception, reasoning, planning, and tool use. However, digital agent systems must often operate in new environments whose interfaces, tools, conventions, and failure modes are not fully captured by pretrained knowledge.

Existing adaptation approaches typically rely on collecting additional interaction data for further training, often with human assistance. This paradigm introduces substantial cost and is difficult to apply in private or continuously changing environments. Training-free adaptation through context management offers a more flexible alternative.

Core Research Question

Can an agent autonomously discover causal relations from a new environment into a reusable memory to improve itself?

The authors draw inspiration from human learning of new software, which follows a recursive loop:

  1. Identifying what needs to be learned
  2. Interacting with the environment to collect experience
  3. Distilling useful knowledge from observed outcomes

This process is framed as a form of causal discovery: by actively trying different actions and observing consequences, agents infer which factors determine success, failure, and state transitions.

Formal Problem Definition

Given a task instruction qq, a digital agent interacts with an environment E\mathcal{E} through a sequence of actions a1,…,aTa_1, \ldots, a_T. At step tt, the agent πθ\pi_\theta generates an executable action ata_t conditioned on durable memory MM and interaction history ht=(o0,a1,…,at−1,ot−1)h_t = (o_0, a_1, \ldots, a_{t-1}, o_{t-1}):

at∼πθ(⋅∣q,ht,M),(st,ot)∼E(⋅∣st−1,at)a_t \sim \pi_\theta(\cdot \mid q, h_t, M), \quad (s_t, o_t) \sim \mathcal{E}(\cdot \mid s_{t-1}, a_t)

where sts_t is the underlying environment state (possibly unobservable), and oto_t is the resulting observation. The interaction history resets between tasks, while durable memory MM persists and retains reusable knowledge across tasks.

The central problem is designing an exploration strategy that:

  • Efficiently identifies informative experiences
  • Grounds them with reliable environment feedback
  • Continuously consolidates resulting knowledge into memory

Methodology

Multi-Agent Harness Framework

RSIAgent decomposes the agent system into three collaborative roles:

1. Actor Agent with Evolvable Memory

  • Primary policy model responsible for understanding the environment and generating executable actions (code-as-policy)
  • Equipped with persistent memory storing environment-specific knowledge, reusable procedures/scripts, and lessons learned
  • Memory is evolvable: after verification, the actor consolidates grounded experience by adding new knowledge, revising or removing outdated information

2. Verifier Agent with Environment Feedback

  • Independent evaluator determining whether execution satisfied task requirements
  • Directly inspects environment feedback (execution results, interface states, observable evidence)
  • Isolated from actor's private reasoning and memory to reduce correlated errors
  • Returns success/failure judgment with supporting feedback

3. Curriculum Agent for Guiding Exploration

  • High-level coordinator deciding what to explore next
  • Generates practice tasks based on current target, accumulated memory, and previous outcomes
  • Selects prerequisite skills, informative variants, failure-driven practice, and stress-test cases
  • Exploration continues until generated tasks are unlikely to contribute substantial new knowledge

Two-Stage Autonomous Exploration

Stage 1: Broad Recursive Self-exploration (BRS)

  • Aims to rapidly build broad understanding of a new environment
  • Follows a recursive exploration loop: curriculum agent generates tasks → actor agents execute in parallel → verifier agents evaluate outcomes
  • At each iteration, proposes multiple tasks spanning different exploration directions executed and verified in parallel
  • Uses accumulated experience to identify knowledge gaps and generate more informative tasks for next iteration

Stage 2: Deep Recursive Self-exploration (DRS)

  • Refines accumulated memory by focusing on important knowledge gaps, hard cases, and boundary conditions
  • Follows a sequential recursive loop that progressively increases difficulty
  • Curriculum agent proposes challenging tasks likely to expose unpredictable issues or hidden constraints
  • Each verified experience is consolidated into memory before the next task is proposed
  • Continuously pushes toward harder and less explored cases

Test-time Memory Reuse

  • After exploration, memory is frozen and provided to the actor agent
  • Curriculum agent and all memory updates are disabled
  • Actor directly reuses procedures, discovered constraints, and failure lessons
  • Action–verification loop continues until the verifier confirms all task requirements are satisfied

Empirical Validation / Results

Experimental Setup

  • Benchmarks: OSWorld 2.0 (0808 offline, 82 tasks) and Agents' Last Exam (ALE) Near-term (67 tasks)
  • Metrics: Partial score (mean task score) and Binary accuracy (proportion of tasks receiving full credit), both as percentages
  • Configuration: GLM-5.3 as default actor, Kimi-K3 as verifier and curriculum agent; BRS budget of 8 exploration projects (up to 4 concurrent); DRS proceeds sequentially until curriculum agent determines no further useful practice

Main Results

Table 1. Model comparison on OSWorld 2.0 (0808 offline) and Agents' Last Exam Near-term

Model / methodOSWorld Partial (%)OSWorld Binary (%)ALE Partial (%)ALE Binary (%)
Open-source Models
Kimi-K2.622.104.6021.709.20
DeepSeek V4 Pro——43.8119.90
Kimi-K358.30—71.6040.30
Closed-source Models
Claude Opus 4.854.8020.6064.0043.30
GPT-5.6 Sol64.1328.1078.8247.76
Claude Opus 570.1934.7279.5446.27
GPT-6 Astra72.60—82.2652.24
Ours (Open-source Models)
RSIAgent (w/o RSI)71.9737.8083.7549.25
RSIAgent78.9842.6884.8250.75

Key findings:

  • RSI improves OSWorld partial score from 71.97 → 78.98 (+7.01) and binary accuracy from 37.80 → 42.68 (+4.88)
  • On ALE, partial score improves from 83.75 → 84.82 and binary accuracy from 49.25 → 50.75
  • RSIAgent exceeds GPT-6 Astra by 6.38 (OSWorld) and 2.56 (ALE) percentage points

Effect of RSI Rounds

Three representative OSWorld 2.0 tasks (T044 video editing, T049 presentation repair, T065 railway booking) show progressive improvement across RSI steps 0–8:

  • By step 8: T044 reaches 100%, T049 reaches 80%, T065 reaches 100%
  • BRS progressively accumulates diverse procedures; DRS refines task-specific details
  • Score increases can be discrete when a remaining bottleneck is resolved

Ablation Study

On four OSWorld 2.0 tasks (T080, T085, T089, T106):

VariantMean Partial Score (%)
w/o RSI (empty memory)~54.06
w/o BRS (deep-only)56.50
w/o DRS (broad-only)65.52
Full RSI74.54
  • Full RSI achieves the highest score on all four tasks
  • Broad-only improves over baseline on every task
  • Deep-only falls below baseline on T085 and T089, highlighting the importance of combining both stages

Game Environment Evaluation (GameCraft-Bench, 40 tasks)

Table 2 highlights (Overall score):

GeneratorBaseline+Play2Code+RSIAgent (w/o RSI)+RSIAgent
Codex + GPT-5.5 (high)52.7751.0557.8461.28
Kimi-K2.631.2836.0242.6146.37
GLM-5.3-Flash30.5538.2544.7348.72
Qwen3.8-27B41.3047.6753.8257.46
  • RSIAgent consistently improves both weak and strong base games
  • Play2Code can degrade high-quality games (51.05 vs 52.77 baseline), while RSIAgent always improves

Failure Mode Analysis

Three mechanisms limiting recursive self-improvement were identified:

  1. Insufficiently Targeted Exploration (75% target-skill mismatch, 25% unchallenged assumptions): Practice may not challenge the decisions responsible for target-specific weaknesses
  2. Incomplete Verification (33.3% grounding gaps, 16.7% coverage gaps, 50% fidelity gaps): Local verification can approve execution without establishing all task requirements satisfied
  3. Unreliable Memory Consolidation (66.7% rule scope loss, 33.3% uncertainty not enforced): Inadequately verified decisions become reusable rules that negatively influence subsequent execution

Theoretical and Practical Implications

Theoretical Contributions

  1. Agent-level self-improvement without parameter updates: Demonstrates that effective adaptation can be achieved purely through memory construction and context management, challenging the assumption that model fine-tuning is necessary for domain adaptation.

  2. Causal discovery through autonomous exploration: Frames the exploration process as a form of active causal discovery, where agents infer stable relationships between actions, conditions, and outcomes through structured experimentation.

  3. Broad-then-deep exploration paradigm: Introduces a coarse-to-fine strategy analogous to the pretraining-then-posttraining paradigm in modern LLM development, providing a principled approach to balancing exploration breadth and depth.

Practical Implications

  1. Bridging the open/closed-source gap: RSIAgent enables open-source models to outperform frontier closed-source models, potentially democratizing access to high-performance agent systems.

  2. Privacy and cost benefits: Training-free adaptation is applicable in private or continuously changing environments where data collection for training is impractical.

  3. Reusable memory: The frozen memory can be directly reused for downstream tasks without additional computation, making the approach practical for deployment.


Conclusion

RSIAgent introduces a training-free multi-agent framework for recursive self-improvement in new digital environments. By coordinating curriculum, actor, and verifier agents with a broad-then-deep exploration strategy, it autonomously acquires, verifies, and consolidates environment-specific knowledge into reusable memory.

Key takeaways:

  • BRS builds diverse environment coverage through parallel exploration
  • DRS refines hard cases, hidden constraints, and boundary conditions through sequential deepening
  • The accumulated memory is frozen and directly reused for downstream tasks

Future directions:

  • Extending beyond digital computer-use environments to broader interactive domains
  • AI for Science applications (specialized tools, workflows, scientific procedures)
  • Games as complex, continuously evolving environments for long-horizon exploration and adaptation

Limitations:

  • Additional test-time exploration introduces substantial computation cost
  • Performance depends on finite exploration budgets, stopping policies, and memory quality
  • Model-based verifier may produce incorrect judgments that propagate
  • Ethical concerns: autonomous exploration risks unintended actions, unauthorized access, and privacy leakage (mitigated by controlled experimental environments)

Related papers