# AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

> AREX-2 shows that training a 27B LLM on long-horizon reflective trajectories in verifiable domains transfers to deep research, achieving state-of-the-art results without new domain data.

- **Source:** [arXiv](https://arxiv.org/abs/2609.38288)
- **Published:** 2026-10-02
- **Permalink:** https://picx.dev/p/iVPsJv
- **Whiteboard:** https://picx.dev/p/iVPsJv/image

## Summary

## Summary (Overview)

- **AREX-2** is a self-improving LLM agent built on Qwen3.8-27B (27B parameters) that achieves state-of-the-art results by learning *long-horizon reflection* — the ability to iteratively refine solutions over many rounds at test time.
- The paper's core hypothesis is that **reflection** (producing better solutions from feedback) and **long-horizon execution** (sustaining improvement over many rounds) are **domain-agnostic meta-skills** that can be learned in verifiable domains (machine learning engineering and algorithmic programming) and transferred to other domains like deep research.
- AREX-2 achieves **81.8 on MLE-bench Lite** (highest among all compared systems) and **70.7 on Frontier-CS** (highest among open-weight models), while also reaching **84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA** — all without adding any new deep-research training data.
- The model demonstrates **test-time scaling**: it keeps improving as its round budget grows, gaining 4.8 points on Frontier-CS between hours 2 and 5, and improving from 64.8 to 84.0 on BrowseComp as turns increase from 47 to 143 — even without external correctness feedback.
- A stage-wise ablation shows that **operational knowledge (skills), training on long-horizon trajectories, and larger round budgets each contribute independently** to the final performance.

---

## Introduction and Theoretical Foundation

### The Problem: Moving the Loop Inside the Model

Hard problems are rarely solved in one attempt. Real-world systems (training pipelines, programs under scoring functions, research questions) improve through a loop of *trying, measuring, and revising*. However, in most agentic systems, this loop is implemented **in the scaffold (harness)** rather than **in the model** — the harness decides when to retry and what to keep, while the model only produces one attempt at a time.

The paper defines **self-improvement** as: *the ability of an agent, given more rounds on a task, to turn them into a better solution by its own judgment of what to change.* This is test-time scaling within a single task.

### Formalization

Let $s_t$ be the best score the agent has reached after $t$ rounds and $r_t = \mathbb{E}[s_t - s_{t-1}]$ the expected gain of round $t$. From a fixed initial score $s_0$ and a budget of $T$ rounds, the expected total improvement is:

$$
\mathbb{E}[s_T] - s_0 = \sum_{t=1}^{T} r_t
\tag{1}
$$

Let $T^* \leq T$ be the number of rounds in which $r_t$ remains non-negligible, and $\bar{r}$ the mean gain over those rounds. The total improvement is approximately $\bar{r} T^*$. This decomposition reveals two complementary capabilities:

- **Reflection** determines $\bar{r}$ — how much a productive round is worth.
- **Long-horizon execution** determines $T^*$ — how many rounds stay productive before the agent stalls.

A self-improving agent needs both, referred to together as **long-horizon reflection**.

### Why Current Training Data Fails

Current agent training data is built *one attempt at a time*: a task is posed, the model produces a solution, a verifier marks pass/fail, and passing solutions are kept. This shows the model what a correct solution looks like, but **not how a solution is improved**. Intermediate attempts, feedback, and revisions are discarded. Long improvement trajectories are the first to be lost.

### The Hypothesis and Domain Selection

> **Hypothesis:** Long-horizon reflection is a meta-skill not tied to any one domain. Judging where a solution falls short, deciding what to change, and continuing after a failed round are the same acts whether the solution is a training pipeline, a program, or a research report.

Two domains are chosen for constructing training trajectories:

1. **Machine learning engineering** (from GitHub repositories)
2. **Algorithmic programming** (from online judges)

These are chosen for three reasons:
- **Unambiguous feedback**: validation scores or judge verdicts directly show whether a revision helped.
- **Room for sustained improvement**: solutions can keep improving over many rounds.
- **Abundant source material**: GitHub and online judges offer large supplies of problems, code, and tests.

---

## Methodology

### 2.1 Self-Improvement as Test-Time Scaling

An **environment** is a pair $e = (\sigma, S)$, where $\sigma$ is the task and $S$ is a scoring function mapping candidate solution $y$ to scalar $S(y)$. In round $t$, the agent submits solution $y_t$ and receives feedback $f_t$ (score, logs, errors, timings). A trajectory of $n$ rounds is:

$$
\tau = \langle y_0, f_0, y_1, f_1, \ldots, y_n, f_n \rangle, \quad s_t = \max_{k \leq t} S(y_k)
\tag{2}
$$

Under a policy $\pi_\theta$, the expected gain of round $t$ is $r_t = \mathbb{E}_{\pi_\theta}[s_t - s_{t-1}]$. The budget $T$ is the resource scaled at test time. For a small threshold $\epsilon$:

$$
T^* = \{t \leq T : r_t > \epsilon\}, \quad \bar{r} = \frac{1}{T^*} \sum_{t \leq T: r_t > \epsilon} r_t
\tag{5}
$$

### 2.2 Constructing Environments

A **teacher model** converts raw sources into executable environments:

- **For GitHub repositories**: The teacher reads the code, identifies the metric being measured, writes a task asking the agent to improve this metric, creates a scoring script computing the metric on a held-out split, and sets up a sandbox with the repository installed.
- **For judge problems**: The task is the problem statement; the sandbox contains a compiler and sample cases; $S$ is the score from hidden tests (graded, not binary, wherever possible).

**Admission criteria** (both checked by execution):
1. The **reference solution** must run under $S$ and obtain a score (environment works).
2. A **simple baseline** must score well below the reference (room for improvement exists).

### 2.3 Constructing Trajectories

Three conditions elicit long-horizon reflection:

1. **More rounds**: The agent gets a budget of rounds and wall-clock time on the order of *hours and hundreds of tool calls per task*, and is told the budget. This changes behavior: establishing a baseline first, using early rounds to measure, later rounds for revisions.

2. **Operational knowledge**: The agent must acquire knowledge about how the environment works (APIs, data loaders, permitted optimizations) by reading source code, searching for papers/libraries, and running small experiments. This knowledge can be provided as **skills** — compact documents placed in the agent's context. *General skills* are prepared in advance; *task-specific skills* are built by searching for task-related information (excluding the task itself).

3. **Feedback in the loop**: Each round ends with a submission; the agent sees the score and environment report before deciding next steps. **Rounds that lower the score are NOT removed** — each is followed by the agent's response, which is what the model should learn.

### 2.4 Training

**Trajectory selection** (whole-trajectory filtering):
- Final score must be high (threshold set relative to reference).
- Process must satisfy task requirements (proper format, every tool call has observation, terminates well-formed, no rule violations).
- **No condition on individual rounds** — failed runs, regressions, and abandoned approaches are deliberately kept to teach recovery from setbacks.

**Supervision within a trajectory**:
- Entire trajectory kept as context; loss applied only to decisions that move the solution forward (diagnosing failure, repairing, changing strategy, running experiments, submitting improvements).
- Steps that make no progress (repeated polling, calls with no new observation, near-duplicate turns) receive no loss.
- System messages, environment reports, and retrieved documents receive no loss.

**Data and model**: Selected trajectories are combined with the unchanged deep-research data from the AREX recipe, and Qwen3.8-27B is fine-tuned to obtain AREX-2. The new trajectories are the **only difference** from the previous recipe, making cross-domain results a clean test of transfer.

---

## Empirical Validation / Results

### 3.1 Benchmarks and Evaluation Protocols

| Benchmark | Capability | Metric |
|---|---|---|
| Frontier-CS | Algorithmic programming (188-task Agent Track) | Accuracy |
| MLE-bench Lite | ML engineering | Any Medal (mean over 3 seeds) |
| BrowseComp | Deep research | Accuracy |
| HLE | Tool-augmented reasoning | Accuracy (text-only subset) |
| GAIA | Information gathering | Accuracy |
| DeepSearchQA | Multi-step task completion | F1 |

Evaluation: Non-coding benchmarks use max 300 inner turns / 1500 total turns (AREX protocol). Frontier-CS uses 5-hour budget; MLE-Lite uses OpenMLE protocol with 12-hour budget.

### 3.2 Overall Results

#### Table 1: Coding and ML Engineering Benchmarks

| Model | Size | Frontier-CS | MLE-Lite |
|---|---|---|---|
| **Closed-Weight** | | | |
| GPT-5.6 Sol | – | 76.4 | 72.7 |
| Claude Opus 4.8 | – | 74.5 | 63.6 |
| GPT-5.5 | – | 72.1 | 68.2 |
| **Open-Weight** | | | |
| Kimi-K3 | 2.8T | – | 72.7 |
| Naive-N0.5-Flash | 309B | – | 73.7 |
| DeepSeek-V4-Pro | 1.6T | 44.7* | 54.5 |
| Kimi-K2.7-Code | 1T | 54.7* | – |
| Frontis-MA1-35B | 35B | – | 71.2 |
| BigBang-V1 | 35B | – | 59.1 |
| **AREX-2** | **27B** | **70.7** | **81.8** |

*AREX-2 achieves the highest MLE-Lite score in the table (81.8, +8.1 over strongest baseline) and the highest Frontier-CS score among open-weight models (70.7, +16.0 over strongest open-weight baseline).*

#### Table 2: Deep Research and Agentic Reasoning Benchmarks (selected)

| Model | Size | BrowseComp | HLE | GAIA | DeepSearchQA |
|---|---|---|---|---|---|
| Kimi-K3 | 2.8T | 91.2 | 56.0* | – | 95.0 |
| Claude Fable 5 | – | 88.0 | 64.5* | – | 94.2 |
| GPT-5.6 Sol | – | 90.4 | 58.0* | – | – |
| Iris-pro | 397B | 88.6 | 56.4 | – | 92.9 |
| AREX (122B) | 122B | 82.5 | 52.4 | 85.4 | 89.9 |
| XYZ-Aquila-mini | 35B | 78.8 | 51.1 | 97.1 | 89.5 |
| Iris-mini | 35B | 82.2 | 52.3 | – | 86.9 |
| AREX (4B) | 4B | 70.7 | 40.6 | 81.6 | 78.5 |
| **AREX-2** | **27B** | **84.0** | **52.6** | **92.2** | **93.8** |

*AREX-2 surpasses both previous AREX models on all four benchmarks despite using the same deep-research training data, and leads all models ≤40B on BrowseComp, HLE, and DeepSearchQA.*

### 3.3 Detailed Analysis

#### 3.3.1 Round Scaling with Feedback (Frontier-CS)

- AREX-2 improves throughout the full 5-hour budget: 54.4 after 1 hour, 65.9 after 2 hours, **70.7 after 5 hours**.
- Gains 4.8 points between hours 2 and 5 (2.2 in the final hour alone).
- Baselines stall: DeepSeek-V4-Pro stops at 44.7 after 2 hours; DeepSeek-V4-Flash at 39.1 after 3 hours.
- **Interpretation**: AREX-2 has a long effective horizon $T^*$; baselines exhaust theirs within 2–3 hours.

#### 3.3.2 Round Scaling without Feedback (BrowseComp)

- AREX-2 accuracy rises steadily with budget: 64.8 at 47 turns → **84.0 at 143 turns**.
- At ~140 turns, AREX-2 is >12 points ahead of AREX (122B) and achieves AREX's final accuracy with **less than half the turns**.
- Gain per turn is ~3× that of AREX (122B), corresponding to a larger $\bar{r}$.
- **Interpretation**: AREX-2 improves without external correctness signals — gains come from self-directed search, reassessment, and refinement. Since both models share the same deep-research data and the new trajectories contain no research tasks, this is evidence of **cross-domain transfer of long-horizon reflection**.

#### 3.3.3 Stage-wise Ablation (MLE-bench Lite)

| Stage | Model | Skills | Rounds | Score |
|---|---|---|---|---|
| M0 | Base (Qwen3.8-27B) | None | Base | 28.8 |
| M1 | Base | General | Base | 41.2 |
| M2 | Base | General + Task-specific | More | 68.2 |
| M3 | AREX-2 | General + Task-specific | Base | 75.8 |
| M4 | AREX-2 | General + Task-specific | More | **81.8** |

**Three observations:**
1. **Operational knowledge matters**: skills + more rounds raise base model from 28.8 → 68.2 without training.
2. **Training adds to skills**: M4 vs M2 (same skills, same rounds) → +13.6 points; M3 vs M2 (fewer rounds) → +7.6 points.
3. **Trained model converts more rounds into improvement**: M3 → M4 gains 6.0 points, consistent with scaling results.

---

## Theoretical and Practical Implications

### Theoretical Implications

1. **Self-improvement is decomposable**: The formalization $\bar{r} T^*$ provides a clean decomposition of test-time scaling into *reflection quality* (gain per round) and *effective horizon* (productive round count). This gives researchers a principled framework for diagnosing why an agent fails to scale: either its reflections are weak ($\bar{r}$ small) or it stalls early ($T^*$ short).

2. **Long-horizon reflection is a transferable meta-skill**: The fact that AREX-2 improves on deep research benchmarks — where no new training data was added — supports the hypothesis that the *process* of iterative improvement is domain-agnostic. This is a significant theoretical claim: the ability to judge, revise, and persist can be learned in one domain and applied in another.

3. **Trajectory-level selection over step-level filtering**: Keeping failed rounds and regressions in training data is not just tolerated but *deliberate*. This challenges the common practice of filtering to successful steps and suggests that learning from setbacks is essential for robustness.

### Practical Implications

1. **Data construction recipe**: The three-step pipeline (construct environments → run teacher agent over many rounds → select whole trajectories) is a practical, reproducible recipe for generating self-improvement training data from abundant sources (GitHub, online judges).

2. **Efficiency**: AREX-2 achieves frontier-competitive results with 27B parameters, matching or exceeding models 10–100× larger. This suggests that *training data quality* (long-horizon trajectories) may matter more than model scale for agentic capabilities.

3. **Operational knowledge via skills**: Providing skills (compact documents in context) yields large gains (+12.4 points from M0→M1, +27 from M1→M2), demonstrating that explicit knowledge injection complements learned reflection.

4. **Test-time scaling is real**: The model converts additional compute (rounds/turns) into better solutions, with diminishing but persistent returns — validating the investment in longer inference budgets.

---

## Conclusion

AREX-2 demonstrates that **moving the improvement loop inside the model** is achievable through training on long-horizon reflective trajectories. The key takeaways:

1. **Self-improvement = reflection × long-horizon execution**: Both capabilities must be learned; neither emerges from training on finished solutions alone.

2. **Verifiable domains are ideal training grounds**: ML engineering and algorithmic programming provide unambiguous feedback, room for sustained improvement, and abundant source material — properties that make them well-suited for supervising the meta-skill.

3. **Transfer works**: What is learned in these domains carries over to deep research, with AREX-2 surpassing previous AREX models on all four research benchmarks without new research data.

4. **Larger budgets convert to better solutions**: AREX-2 keeps improving through its full budget (5 hours on Frontier-CS) and even without correctness feedback (BrowseComp), with gain-per-turn ~3× that of the previous recipe.

**Future directions** identified by the authors:
- **Widen the training domains** beyond ML engineering and algorithmic programming.
- **Lengthen the horizons** — even longer trajectories may yield further gains.
- **Close the loop** — let the model's own trajectories become its next training data (self-improving training).

The paper concludes that *long-horizon reflection, learned where it can be supervised, carries over beyond the domains it was learned in* — offering a scalable path toward genuinely self-improving agents.

---

_Markdown view of https://picx.dev/p/iVPsJv, served by PicX — AI-generated visual whiteboard summaries of research papers._
