Full text not available for this paper
Summary (Overview)
- AREX-2 is a self-improving LLM agent built on Qwen3.8-27B (27B parameters) that achieves state-of-the-art results by learning long-horizon reflection â the ability to iteratively refine solutions over many rounds at test time.
- The paper's core hypothesis is that reflection (producing better solutions from feedback) and long-horizon execution (sustaining improvement over many rounds) are domain-agnostic meta-skills that can be learned in verifiable domains (machine learning engineering and algorithmic programming) and transferred to other domains like deep research.
- AREX-2 achieves 81.8 on MLE-bench Lite (highest among all compared systems) and 70.7 on Frontier-CS (highest among open-weight models), while also reaching 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA â all without adding any new deep-research training data.
- The model demonstrates test-time scaling: it keeps improving as its round budget grows, gaining 4.8 points on Frontier-CS between hours 2 and 5, and improving from 64.8 to 84.0 on BrowseComp as turns increase from 47 to 143 â even without external correctness feedback.
- A stage-wise ablation shows that operational knowledge (skills), training on long-horizon trajectories, and larger round budgets each contribute independently to the final performance.
Introduction and Theoretical Foundation
The Problem: Moving the Loop Inside the Model
Hard problems are rarely solved in one attempt. Real-world systems (training pipelines, programs under scoring functions, research questions) improve through a loop of trying, measuring, and revising. However, in most agentic systems, this loop is implemented in the scaffold (harness) rather than in the model â the harness decides when to retry and what to keep, while the model only produces one attempt at a time.
The paper defines self-improvement as: the ability of an agent, given more rounds on a task, to turn them into a better solution by its own judgment of what to change. This is test-time scaling within a single task.
Formalization
Let be the best score the agent has reached after rounds and the expected gain of round . From a fixed initial score and a budget of rounds, the expected total improvement is:
Let be the number of rounds in which remains non-negligible, and the mean gain over those rounds. The total improvement is approximately . This decomposition reveals two complementary capabilities:
- Reflection determines â how much a productive round is worth.
- Long-horizon execution determines â how many rounds stay productive before the agent stalls.
A self-improving agent needs both, referred to together as long-horizon reflection.
Why Current Training Data Fails
Current agent training data is built one attempt at a time: a task is posed, the model produces a solution, a verifier marks pass/fail, and passing solutions are kept. This shows the model what a correct solution looks like, but not how a solution is improved. Intermediate attempts, feedback, and revisions are discarded. Long improvement trajectories are the first to be lost.
The Hypothesis and Domain Selection
Hypothesis: Long-horizon reflection is a meta-skill not tied to any one domain. Judging where a solution falls short, deciding what to change, and continuing after a failed round are the same acts whether the solution is a training pipeline, a program, or a research report.
Two domains are chosen for constructing training trajectories:
- Machine learning engineering (from GitHub repositories)
- Algorithmic programming (from online judges)
These are chosen for three reasons:
- Unambiguous feedback: validation scores or judge verdicts directly show whether a revision helped.
- Room for sustained improvement: solutions can keep improving over many rounds.
- Abundant source material: GitHub and online judges offer large supplies of problems, code, and tests.
Methodology
2.1 Self-Improvement as Test-Time Scaling
An environment is a pair , where is the task and is a scoring function mapping candidate solution to scalar . In round , the agent submits solution and receives feedback (score, logs, errors, timings). A trajectory of rounds is:
Under a policy , the expected gain of round is . The budget is the resource scaled at test time. For a small threshold :
2.2 Constructing Environments
A teacher model converts raw sources into executable environments:
- For GitHub repositories: The teacher reads the code, identifies the metric being measured, writes a task asking the agent to improve this metric, creates a scoring script computing the metric on a held-out split, and sets up a sandbox with the repository installed.
- For judge problems: The task is the problem statement; the sandbox contains a compiler and sample cases; is the score from hidden tests (graded, not binary, wherever possible).
Admission criteria (both checked by execution):
- The reference solution must run under and obtain a score (environment works).
- A simple baseline must score well below the reference (room for improvement exists).
2.3 Constructing Trajectories
Three conditions elicit long-horizon reflection:
-
More rounds: The agent gets a budget of rounds and wall-clock time on the order of hours and hundreds of tool calls per task, and is told the budget. This changes behavior: establishing a baseline first, using early rounds to measure, later rounds for revisions.
-
Operational knowledge: The agent must acquire knowledge about how the environment works (APIs, data loaders, permitted optimizations) by reading source code, searching for papers/libraries, and running small experiments. This knowledge can be provided as skills â compact documents placed in the agent's context. General skills are prepared in advance; task-specific skills are built by searching for task-related information (excluding the task itself).
-
Feedback in the loop: Each round ends with a submission; the agent sees the score and environment report before deciding next steps. Rounds that lower the score are NOT removed â each is followed by the agent's response, which is what the model should learn.
2.4 Training
Trajectory selection (whole-trajectory filtering):
- Final score must be high (threshold set relative to reference).
- Process must satisfy task requirements (proper format, every tool call has observation, terminates well-formed, no rule violations).
- No condition on individual rounds â failed runs, regressions, and abandoned approaches are deliberately kept to teach recovery from setbacks.
Supervision within a trajectory:
- Entire trajectory kept as context; loss applied only to decisions that move the solution forward (diagnosing failure, repairing, changing strategy, running experiments, submitting improvements).
- Steps that make no progress (repeated polling, calls with no new observation, near-duplicate turns) receive no loss.
- System messages, environment reports, and retrieved documents receive no loss.
Data and model: Selected trajectories are combined with the unchanged deep-research data from the AREX recipe, and Qwen3.8-27B is fine-tuned to obtain AREX-2. The new trajectories are the only difference from the previous recipe, making cross-domain results a clean test of transfer.
Empirical Validation / Results
3.1 Benchmarks and Evaluation Protocols
| Benchmark | Capability | Metric |
|---|---|---|
| Frontier-CS | Algorithmic programming (188-task Agent Track) | Accuracy |
| MLE-bench Lite | ML engineering | Any Medal (mean over 3 seeds) |
| BrowseComp | Deep research | Accuracy |
| HLE | Tool-augmented reasoning | Accuracy (text-only subset) |
| GAIA | Information gathering | Accuracy |
| DeepSearchQA | Multi-step task completion | F1 |
Evaluation: Non-coding benchmarks use max 300 inner turns / 1500 total turns (AREX protocol). Frontier-CS uses 5-hour budget; MLE-Lite uses OpenMLE protocol with 12-hour budget.
3.2 Overall Results
Table 1: Coding and ML Engineering Benchmarks
| Model | Size | Frontier-CS | MLE-Lite |
|---|---|---|---|
| Closed-Weight | |||
| GPT-5.6 Sol | â | 76.4 | 72.7 |
| Claude Opus 4.8 | â | 74.5 | 63.6 |
| GPT-5.5 | â | 72.1 | 68.2 |
| Open-Weight | |||
| Kimi-K3 | 2.8T | â | 72.7 |
| Naive-N0.5-Flash | 309B | â | 73.7 |
| DeepSeek-V4-Pro | 1.6T | 44.7* | 54.5 |
| Kimi-K2.7-Code | 1T | 54.7* | â |
| Frontis-MA1-35B | 35B | â | 71.2 |
| BigBang-V1 | 35B | â | 59.1 |
| AREX-2 | 27B | 70.7 | 81.8 |
AREX-2 achieves the highest MLE-Lite score in the table (81.8, +8.1 over strongest baseline) and the highest Frontier-CS score among open-weight models (70.7, +16.0 over strongest open-weight baseline).
Table 2: Deep Research and Agentic Reasoning Benchmarks (selected)
| Model | Size | BrowseComp | HLE | GAIA | DeepSearchQA |
|---|---|---|---|---|---|
| Kimi-K3 | 2.8T | 91.2 | 56.0* | â | 95.0 |
| Claude Fable 5 | â | 88.0 | 64.5* | â | 94.2 |
| GPT-5.6 Sol | â | 90.4 | 58.0* | â | â |
| Iris-pro | 397B | 88.6 | 56.4 | â | 92.9 |
| AREX (122B) | 122B | 82.5 | 52.4 | 85.4 | 89.9 |
| XYZ-Aquila-mini | 35B | 78.8 | 51.1 | 97.1 | 89.5 |
| Iris-mini | 35B | 82.2 | 52.3 | â | 86.9 |
| AREX (4B) | 4B | 70.7 | 40.6 | 81.6 | 78.5 |
| AREX-2 | 27B | 84.0 | 52.6 | 92.2 | 93.8 |
AREX-2 surpasses both previous AREX models on all four benchmarks despite using the same deep-research training data, and leads all models â€40B on BrowseComp, HLE, and DeepSearchQA.
3.3 Detailed Analysis
3.3.1 Round Scaling with Feedback (Frontier-CS)
- AREX-2 improves throughout the full 5-hour budget: 54.4 after 1 hour, 65.9 after 2 hours, 70.7 after 5 hours.
- Gains 4.8 points between hours 2 and 5 (2.2 in the final hour alone).
- Baselines stall: DeepSeek-V4-Pro stops at 44.7 after 2 hours; DeepSeek-V4-Flash at 39.1 after 3 hours.
- Interpretation: AREX-2 has a long effective horizon ; baselines exhaust theirs within 2â3 hours.
3.3.2 Round Scaling without Feedback (BrowseComp)
- AREX-2 accuracy rises steadily with budget: 64.8 at 47 turns â 84.0 at 143 turns.
- At ~140 turns, AREX-2 is >12 points ahead of AREX (122B) and achieves AREX's final accuracy with less than half the turns.
- Gain per turn is ~3Ã that of AREX (122B), corresponding to a larger .
- Interpretation: AREX-2 improves without external correctness signals â gains come from self-directed search, reassessment, and refinement. Since both models share the same deep-research data and the new trajectories contain no research tasks, this is evidence of cross-domain transfer of long-horizon reflection.
3.3.3 Stage-wise Ablation (MLE-bench Lite)
| Stage | Model | Skills | Rounds | Score |
|---|---|---|---|---|
| M0 | Base (Qwen3.8-27B) | None | Base | 28.8 |
| M1 | Base | General | Base | 41.2 |
| M2 | Base | General + Task-specific | More | 68.2 |
| M3 | AREX-2 | General + Task-specific | Base | 75.8 |
| M4 | AREX-2 | General + Task-specific | More | 81.8 |
Three observations:
- Operational knowledge matters: skills + more rounds raise base model from 28.8 â 68.2 without training.
- Training adds to skills: M4 vs M2 (same skills, same rounds) â +13.6 points; M3 vs M2 (fewer rounds) â +7.6 points.
- Trained model converts more rounds into improvement: M3 â M4 gains 6.0 points, consistent with scaling results.
Theoretical and Practical Implications
Theoretical Implications
-
Self-improvement is decomposable: The formalization provides a clean decomposition of test-time scaling into reflection quality (gain per round) and effective horizon (productive round count). This gives researchers a principled framework for diagnosing why an agent fails to scale: either its reflections are weak ( small) or it stalls early ( short).
-
Long-horizon reflection is a transferable meta-skill: The fact that AREX-2 improves on deep research benchmarks â where no new training data was added â supports the hypothesis that the process of iterative improvement is domain-agnostic. This is a significant theoretical claim: the ability to judge, revise, and persist can be learned in one domain and applied in another.
-
Trajectory-level selection over step-level filtering: Keeping failed rounds and regressions in training data is not just tolerated but deliberate. This challenges the common practice of filtering to successful steps and suggests that learning from setbacks is essential for robustness.
Practical Implications
-
Data construction recipe: The three-step pipeline (construct environments â run teacher agent over many rounds â select whole trajectories) is a practical, reproducible recipe for generating self-improvement training data from abundant sources (GitHub, online judges).
-
Efficiency: AREX-2 achieves frontier-competitive results with 27B parameters, matching or exceeding models 10â100Ã larger. This suggests that training data quality (long-horizon trajectories) may matter more than model scale for agentic capabilities.
-
Operational knowledge via skills: Providing skills (compact documents in context) yields large gains (+12.4 points from M0âM1, +27 from M1âM2), demonstrating that explicit knowledge injection complements learned reflection.
-
Test-time scaling is real: The model converts additional compute (rounds/turns) into better solutions, with diminishing but persistent returns â validating the investment in longer inference budgets.
Conclusion
AREX-2 demonstrates that moving the improvement loop inside the model is achievable through training on long-horizon reflective trajectories. The key takeaways:
-
Self-improvement = reflection à long-horizon execution: Both capabilities must be learned; neither emerges from training on finished solutions alone.
-
Verifiable domains are ideal training grounds: ML engineering and algorithmic programming provide unambiguous feedback, room for sustained improvement, and abundant source material â properties that make them well-suited for supervising the meta-skill.
-
Transfer works: What is learned in these domains carries over to deep research, with AREX-2 surpassing previous AREX models on all four research benchmarks without new research data.
-
Larger budgets convert to better solutions: AREX-2 keeps improving through its full budget (5 hours on Frontier-CS) and even without correctness feedback (BrowseComp), with gain-per-turn ~3Ã that of the previous recipe.
Future directions identified by the authors:
- Widen the training domains beyond ML engineering and algorithmic programming.
- Lengthen the horizons â even longer trajectories may yield further gains.
- Close the loop â let the model's own trajectories become its next training data (self-improving training).
The paper concludes that long-horizon reflection, learned where it can be supervised, carries over beyond the domains it was learned in â offering a scalable path toward genuinely self-improving agents.
Related papers
- Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
ReaLVR closes the latent evidence-credit gap by supervising free-running latent tokens with answer-contrastive readouts and visual prototypes, achieving 63.7% average accuracy and scaling to 235B parameters.
- The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation
RIDE extrapolates RL-induced hidden-state residuals beyond the teacher, consistently surpassing it across four model pairs where output-space extrapolation fails.
- Post-Training Leaves Behavioral Shadows on Unrelated Decisions
Post-training updates leave behavioral shadows on unrelated inputs that can be extracted via single-word queries to transfer capabilities, yielding +5.34 points on HumanEval+.