# RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments

> RSIAgent, a training-free multi-agent framework, enables open-source models to outperform frontier closed-source models by autonomously exploring, verifying, and consolidating environment knowledge into reusable memory.

- **Source:** [arXiv](https://arxiv.org/abs/2609.15364)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/hWwdMO
- **Whiteboard:** https://picx.dev/p/hWwdMO/image

## Summary

## Summary (Overview)

- **RSIAgent** is a training-free multi-agent framework for recursive self-improvement in new digital environments, enabling agents to autonomously explore, verify, and consolidate environment-specific knowledge into reusable memory without updating model parameters.
- The framework coordinates three agent roles: a **curriculum agent** (decides what to explore), an **actor agent** (executes actions and updates memory), and a **verifier agent** (grounds outcomes in environment feedback).
- A **broad-then-deep exploration strategy** is introduced: *Broad Recursive Self-exploration (BRS)* builds diverse environment coverage in parallel, while *Deep Recursive Self-exploration (DRS)* focuses on hard cases, hidden constraints, and boundary conditions.
- Experiments on **OSWorld-v2** and **Agent's Last Exam (ALE)** show that RSIAgent enables open-source models (Kimi-K3, GLM-5.3) to outperform frontier closed-source models including GPT-6 Astra and Claude Opus 5.
- Additional evaluation on **GameCraft-Bench** demonstrates generalization beyond computer-use tasks to interactive game development, consistently improving game quality across both weak and strong base generators.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Large language models (LLMs) and vision-language models (VLMs) have advanced significantly in perception, reasoning, planning, and tool use. However, digital agent systems must often operate in **new environments** whose interfaces, tools, conventions, and failure modes are not fully captured by pretrained knowledge.

Existing adaptation approaches typically rely on collecting additional interaction data for further training, often with human assistance. This paradigm introduces substantial cost and is difficult to apply in private or continuously changing environments. **Training-free adaptation** through context management offers a more flexible alternative.

### Core Research Question

> Can an agent autonomously discover causal relations from a new environment into a reusable memory to improve itself?

The authors draw inspiration from human learning of new software, which follows a recursive loop:
1. Identifying what needs to be learned
2. Interacting with the environment to collect experience
3. Distilling useful knowledge from observed outcomes

This process is framed as a form of **causal discovery**: by actively trying different actions and observing consequences, agents infer which factors determine success, failure, and state transitions.

### Formal Problem Definition

Given a task instruction $q$, a digital agent interacts with an environment $\mathcal{E}$ through a sequence of actions $a_1, \ldots, a_T$. At step $t$, the agent $\pi_\theta$ generates an executable action $a_t$ conditioned on durable memory $M$ and interaction history $h_t = (o_0, a_1, \ldots, a_{t-1}, o_{t-1})$:

$$a_t \sim \pi_\theta(\cdot \mid q, h_t, M), \quad (s_t, o_t) \sim \mathcal{E}(\cdot \mid s_{t-1}, a_t)$$

where $s_t$ is the underlying environment state (possibly unobservable), and $o_t$ is the resulting observation. The interaction history resets between tasks, while **durable memory** $M$ persists and retains reusable knowledge across tasks.

The central problem is designing an exploration strategy that:
- Efficiently identifies informative experiences
- Grounds them with reliable environment feedback
- Continuously consolidates resulting knowledge into memory

---

## Methodology

### Multi-Agent Harness Framework

RSIAgent decomposes the agent system into three collaborative roles:

#### 1. Actor Agent with Evolvable Memory
- Primary policy model responsible for understanding the environment and generating executable actions (code-as-policy)
- Equipped with persistent memory storing environment-specific knowledge, reusable procedures/scripts, and lessons learned
- Memory is **evolvable**: after verification, the actor consolidates grounded experience by adding new knowledge, revising or removing outdated information

#### 2. Verifier Agent with Environment Feedback
- Independent evaluator determining whether execution satisfied task requirements
- Directly inspects environment feedback (execution results, interface states, observable evidence)
- **Isolated** from actor's private reasoning and memory to reduce correlated errors
- Returns success/failure judgment with supporting feedback

#### 3. Curriculum Agent for Guiding Exploration
- High-level coordinator deciding what to explore next
- Generates practice tasks based on current target, accumulated memory, and previous outcomes
- Selects prerequisite skills, informative variants, failure-driven practice, and stress-test cases
- Exploration continues until generated tasks are unlikely to contribute substantial new knowledge

### Two-Stage Autonomous Exploration

#### Stage 1: Broad Recursive Self-exploration (BRS)
- Aims to rapidly build broad understanding of a new environment
- Follows a **recursive exploration loop**: curriculum agent generates tasks → actor agents execute in parallel → verifier agents evaluate outcomes
- At each iteration, proposes **multiple tasks spanning different exploration directions** executed and verified in parallel
- Uses accumulated experience to identify knowledge gaps and generate more informative tasks for next iteration

#### Stage 2: Deep Recursive Self-exploration (DRS)
- Refines accumulated memory by focusing on important knowledge gaps, hard cases, and boundary conditions
- Follows a **sequential recursive loop** that progressively increases difficulty
- Curriculum agent proposes challenging tasks likely to expose unpredictable issues or hidden constraints
- Each verified experience is consolidated into memory **before** the next task is proposed
- Continuously pushes toward harder and less explored cases

#### Test-time Memory Reuse
- After exploration, memory is **frozen** and provided to the actor agent
- Curriculum agent and all memory updates are **disabled**
- Actor directly reuses procedures, discovered constraints, and failure lessons
- Action–verification loop continues until the verifier confirms all task requirements are satisfied

---

## Empirical Validation / Results

### Experimental Setup

- **Benchmarks**: OSWorld 2.0 (0808 offline, 82 tasks) and Agents' Last Exam (ALE) Near-term (67 tasks)
- **Metrics**: Partial score (mean task score) and Binary accuracy (proportion of tasks receiving full credit), both as percentages
- **Configuration**: GLM-5.3 as default actor, Kimi-K3 as verifier and curriculum agent; BRS budget of 8 exploration projects (up to 4 concurrent); DRS proceeds sequentially until curriculum agent determines no further useful practice

### Main Results

**Table 1. Model comparison on OSWorld 2.0 (0808 offline) and Agents' Last Exam Near-term**

| Model / method | OSWorld Partial (%) | OSWorld Binary (%) | ALE Partial (%) | ALE Binary (%) |
|---|---|---|---|---|
| **Open-source Models** | | | | |
| Kimi-K2.6 | 22.10 | 4.60 | 21.70 | 9.20 |
| DeepSeek V4 Pro | — | — | 43.81 | 19.90 |
| Kimi-K3 | 58.30 | — | 71.60 | 40.30 |
| **Closed-source Models** | | | | |
| Claude Opus 4.8 | 54.80 | 20.60 | 64.00 | 43.30 |
| GPT-5.6 Sol | 64.13 | 28.10 | 78.82 | 47.76 |
| Claude Opus 5 | 70.19 | 34.72 | 79.54 | 46.27 |
| GPT-6 Astra | 72.60 | — | 82.26 | 52.24 |
| **Ours (Open-source Models)** | | | | |
| RSIAgent (w/o RSI) | 71.97 | 37.80 | 83.75 | 49.25 |
| **RSIAgent** | **78.98** | **42.68** | **84.82** | **50.75** |

Key findings:
- RSI improves OSWorld partial score from 71.97 → **78.98** (+7.01) and binary accuracy from 37.80 → **42.68** (+4.88)
- On ALE, partial score improves from 83.75 → **84.82** and binary accuracy from 49.25 → **50.75**
- RSIAgent exceeds GPT-6 Astra by **6.38** (OSWorld) and **2.56** (ALE) percentage points

### Effect of RSI Rounds

Three representative OSWorld 2.0 tasks (T044 video editing, T049 presentation repair, T065 railway booking) show progressive improvement across RSI steps 0–8:
- By step 8: T044 reaches **100%**, T049 reaches **80%**, T065 reaches **100%**
- BRS progressively accumulates diverse procedures; DRS refines task-specific details
- Score increases can be **discrete** when a remaining bottleneck is resolved

### Ablation Study

On four OSWorld 2.0 tasks (T080, T085, T089, T106):

| Variant | Mean Partial Score (%) |
|---|---|
| w/o RSI (empty memory) | ~54.06 |
| w/o BRS (deep-only) | 56.50 |
| w/o DRS (broad-only) | 65.52 |
| **Full RSI** | **74.54** |

- Full RSI achieves the highest score on **all four tasks**
- Broad-only improves over baseline on every task
- Deep-only falls below baseline on T085 and T089, highlighting the importance of combining both stages

### Game Environment Evaluation (GameCraft-Bench, 40 tasks)

**Table 2 highlights (Overall score):**

| Generator | Baseline | +Play2Code | +RSIAgent (w/o RSI) | +RSIAgent |
|---|---|---|---|---|
| Codex + GPT-5.5 (high) | 52.77 | 51.05 | 57.84 | **61.28** |
| Kimi-K2.6 | 31.28 | 36.02 | 42.61 | **46.37** |
| GLM-5.3-Flash | 30.55 | 38.25 | 44.73 | **48.72** |
| Qwen3.8-27B | 41.30 | 47.67 | 53.82 | **57.46** |

- RSIAgent consistently improves both weak and strong base games
- Play2Code can **degrade** high-quality games (51.05 vs 52.77 baseline), while RSIAgent always improves

### Failure Mode Analysis

Three mechanisms limiting recursive self-improvement were identified:

1. **Insufficiently Targeted Exploration** (75% target-skill mismatch, 25% unchallenged assumptions): Practice may not challenge the decisions responsible for target-specific weaknesses
2. **Incomplete Verification** (33.3% grounding gaps, 16.7% coverage gaps, 50% fidelity gaps): Local verification can approve execution without establishing all task requirements satisfied
3. **Unreliable Memory Consolidation** (66.7% rule scope loss, 33.3% uncertainty not enforced): Inadequately verified decisions become reusable rules that negatively influence subsequent execution

---

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Agent-level self-improvement without parameter updates**: Demonstrates that effective adaptation can be achieved purely through memory construction and context management, challenging the assumption that model fine-tuning is necessary for domain adaptation.

2. **Causal discovery through autonomous exploration**: Frames the exploration process as a form of active causal discovery, where agents infer stable relationships between actions, conditions, and outcomes through structured experimentation.

3. **Broad-then-deep exploration paradigm**: Introduces a coarse-to-fine strategy analogous to the pretraining-then-posttraining paradigm in modern LLM development, providing a principled approach to balancing exploration breadth and depth.

### Practical Implications

1. **Bridging the open/closed-source gap**: RSIAgent enables open-source models to outperform frontier closed-source models, potentially democratizing access to high-performance agent systems.

2. **Privacy and cost benefits**: Training-free adaptation is applicable in private or continuously changing environments where data collection for training is impractical.

3. **Reusable memory**: The frozen memory can be directly reused for downstream tasks without additional computation, making the approach practical for deployment.

---

## Conclusion

RSIAgent introduces a **training-free multi-agent framework** for recursive self-improvement in new digital environments. By coordinating curriculum, actor, and verifier agents with a broad-then-deep exploration strategy, it autonomously acquires, verifies, and consolidates environment-specific knowledge into reusable memory.

**Key takeaways:**
- BRS builds diverse environment coverage through parallel exploration
- DRS refines hard cases, hidden constraints, and boundary conditions through sequential deepening
- The accumulated memory is frozen and directly reused for downstream tasks

**Future directions:**
- Extending beyond digital computer-use environments to broader interactive domains
- AI for Science applications (specialized tools, workflows, scientific procedures)
- Games as complex, continuously evolving environments for long-horizon exploration and adaptation

**Limitations:**
- Additional test-time exploration introduces substantial computation cost
- Performance depends on finite exploration budgets, stopping policies, and memory quality
- Model-based verifier may produce incorrect judgments that propagate
- Ethical concerns: autonomous exploration risks unintended actions, unauthorized access, and privacy leakage (mitigated by controlled experimental environments)

---

_Markdown view of https://picx.dev/p/hWwdMO, served by PicX — AI-generated visual whiteboard summaries of research papers._
