# Stateless Language Agents: Scaling Long-Horizon Automated Research

> Stateless Language Agents, where the harness owns all research state and reconstructs fresh contexts per invocation, outperform stateful agent frameworks on long-horizon tasks, reaching baseline final performance with over 84% fewer tokens.

- **Source:** [arXiv](https://arxiv.org/abs/2610.07625)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/707feb
- **Whiteboard:** https://picx.dev/p/707feb/image

## Summary

## Summary (Overview)

- **Core contribution**: Introduces **Stateless Language Agents (SLAs)**, a framework for long-horizon automated research built on the principle of *stateful search with stateless agents* — the harness owns all research state while each agent invocation starts fresh with a reconstructed, role-specific context.
- **Key finding**: SLA achieves the best final result on every evaluated task (software engineering, kernel optimization, algorithm design) at budgets up to 1 billion tokens, reaching the strongest baseline's final performance with over 84% fewer tokens.
- **Design insight**: Focused contexts and explicit work assignments each contribute to progress, with effects that compound over long horizons; removing Worker context isolation costs 5–11 cycles over 100M-token continuations but 224 cycles over a full 1B-token run.
- **Evaluation contribution**: Demonstrates that short evaluation horizons can misjudge research systems — SLA trails baselines at 25% of the budget on FrontierSWE but leads at 100%.
- **Efficiency**: The Advisor (coordinator) consumes only 0.24–0.51% of tokens and 1.2–2.3% of model cost, compared to 8.4–10.7% of tokens for SwarmResearch's Shepherd.

## Introduction and Theoretical Foundation

### Background and Motivation

Automated research (AutoResearch) systems use LLM agents to iteratively propose, implement, and evaluate solutions to problems with measurable objectives but no known optimum. As these systems tackle harder problems, the critical question becomes: **how can automated research sustain progress as its inference budget grows?**

The paper identifies two fundamental challenges in long-horizon runs:

1. **Retaining experience while controlling context**: Candidate solutions and evaluation results are reusable evidence, but replaying growing histories consumes tokens and degrades agent behavior. Accumulated context can cause premature stopping in long-horizon search.

2. **Turning accumulated evidence into productive exploration**: Recorded findings don't determine what to try next. In one CORAL run, the last improvement occurred at 231M tokens, after which agents repeatedly declared solutions optimal — 98% of the final 1,500 sessions made no tool calls.

### Key Theoretical Distinction

The paper introduces a crucial separation:

> **State** (persists across invocations) vs. **Context** (what an agent sees within one invocation)

A stateless agent keeps no state — nothing carries over between invocations. This does not make it contextless; rather, the harness reconstructs a fresh, role-specific context for every invocation, making what each agent sees an **explicit design choice** rather than a by-product of accumulated conversation.

### Limitations of Existing Evaluations

Most evaluations use short budgets or benchmarks that saturate early. On 26-circle packing, four methods finish within $2.3 \times 10^{-10}$ of one another by 5M tokens — a saturated benchmark measures speed to ceiling, not sustained progress. Additionally, wall-clock time mixes model latency, evaluator runtime, and parallelism, so the paper argues for measuring horizons in **cumulative tokens**.

## Methodology

### The SLA Framework

The SLA framework alternates global planning with local experimentation:

```
Each epoch:
1. Harness reconstructs Advisor's context from research state
2. Advisor assigns concrete experiments to parallel Workers
3. Each Worker runs several local trials in isolated workspace
4. Independent evaluator scores every candidate
5. Harness records outcomes, updates local candidates and global best
```

### Harness-Controlled Context Reconstruction

**Advisor context**: Receives the task specification, current global best code/score, and a harness-built evidence summary covering:
- Recent attempts and outcomes grouped by search direction
- Worker status and recent progress
- Task-specific diagnostics
- Separation of measured non-improvements from inconclusive failures

**Worker context**: Receives only:
- Task specification
- Its assignment
- Its retained candidate
- Compact local feedback (previous trial status, measured score, correctness diagnostics)

Workers cannot access other Workers' workspaces or the full research history.

### Evidence-Driven Work Assignment

The Advisor turns global evidence into concrete, coordinated assignments. Workers decide how to implement and refine assignments over several local trials. When redirected, Workers restart from the global best, with their first successful candidate becoming a new local baseline.

### Implementation Details

- Uses OpenAI Codex or Claude Code for both Advisor and Workers
- Statelessness enforced via new sessions each invocation
- Workers run as unprivileged processes under Linux Landlock filesystem restrictions
- Candidates evaluated outside Worker environments with protected evaluator code
- Full research state checkpoints enable controlled ablations

### Evaluation Setup

**Tasks**: FrontierSWE (4 problems), Anthropic VLIW SIMD kernel optimization, SOL-ExecBench GPU kernels (3 problems), FrontierCS Structured-LWE.

**Baselines**: EvoX, CORAL, SwarmResearch — all from pinned upstream releases with unmodified search logic.

**Models**: OpenAI Codex CLI v0.152.1 with GPT-5.5; Claude Code v2.1.258 with Claude Opus 4.8.

## Empirical Validation / Results

### Main Results: Performance and Token Efficiency

**Table 1: Anthropic VLIW SIMD kernel performance and token efficiency**

| Method | 250M Cycles ↓ | 1B Cycles ↓ | To match Tokens (M) ↓ |
|---|---|---|---|
| EvoX | $1478.3 \pm 112.0$ | $1343.3 \pm 4.0$ | NR |
| CORAL | $1524.0 \pm 101.7$ | $1350.0 \pm 134.2$ | NR |
| SwarmResearch | $1501.7 \pm 185.7$ | $1275.7 \pm 133.9$ | $986.3$ |
| **SLA (ours)** | **$1149.3 \pm 18.9$** | **$1112.0 \pm 8.9$** | **67.9** |
| Reduction vs. best baseline | -22.3% | -12.8% | -93.1% |

Key findings:
- All three SLA runs outperform all nine baseline runs (worst SLA: 1122 cycles vs. best baseline: 1191 cycles)
- SLA reaches the strongest baseline's final performance with 93.1% fewer tokens (Codex) and 84.4% fewer (Claude Code)
- On FrontierSWE, SLA trails at 175M tokens but leads at 700M on all four tasks (average improvement: +4.95 points)
- SOL-ExecBench scores improve by 5.5% on average; FrontierCS Structured-LWE improves by 1.5 points

### Ablation Studies

From shared checkpoints, removing any design component reduces mean progress:

| Configuration | Anthropic Kernel (100M→200M) | FrontierSWE libexpat (70M→170M) |
|---|---|---|
| SLA (full) | $1192.3 \pm 2.5$ (-25.7) | $18.78 \pm 0.92$ (+2.66) |
| w/o Advisor reconstruction | $1198.3 \pm 2.3$ (-23.3%) | $16.86 \pm 0.64$ (-72.2%) |
| w/o Worker isolation | $1203.3 \pm 7.1$ (-42.8%) | $16.79 \pm 0.47$ (-74.8%) |
| w/o Advisor assignments | $1195.0 \pm 2.6$ (-10.5%) | $16.54 \pm 0.33$ (-84.2%) |

**Critical finding**: Removing Worker isolation costs only 5–11 cycles over 100M-token continuations but **224 cycles** over a full 1B-token run (1336 vs. 1112 cycles) — effects compound over longer horizons.

### Scaling and Coordination Cost

- **Coordination is cheap**: Advisor uses 0.24–0.51% of tokens vs. 8.39–10.70% for SwarmResearch's Shepherd
- **Worker scaling**: Going from 1 to 15 Workers cuts elapsed time by 6.6–7.3×; wider pools that trail early can catch up (e.g., W=15 trails by 337 cycles at 25% budget but finishes within 4 cycles at 100%)
- **Task-dependent optimal width**: Dart peaks at W=3, Git/libexpat at W=7, Lua at W=1
- **Worker model choice**: GPT-5.4 mini Workers outperform GPT-5.5 on Git→Zig at matched cost ($100), but GPT-5.5 performs better on kernel optimization

## Theoretical and Practical Implications

### Design Principles for Long-Horizon Research Systems

The paper argues that keeping research state in the harness makes three things explicit design choices:
1. What the system retains
2. What each agent sees
3. Who chooses the next experiment

These choices compound over time, so short evaluations can misjudge both methods and their components.

### Implications for Post-Training

In SLA, a long run becomes a sequence of short invocations:
- Each Worker trial has bounded context and ends with a score — usable directly as RL episodes
- Models don't need training on full billion-token runs
- Advisor assignments with measured outcomes can train experiment-planning models
- Short contexts may reduce dependence on long-context ability

### Implications for Safety

- All persistent state lives in the harness — inspectable, checkpointable, rollback-able
- No agent carries plans or injected instructions forward in its own conversation
- Protected evaluators remain necessary (candidate code could still game the evaluator)

### Evaluation Methodology

The paper advocates reporting results at **multiple cumulative token budgets**, as short horizons can misjudge:
- Methods (SLA trails at 25% budget, leads at 100% on FrontierSWE)
- Components (Worker isolation cost grows >10× between short and full runs)
- Configurations (wider Worker pools that trail early can catch up)

## Conclusion

The paper demonstrates that **stateless language agents** — where the harness owns research state and reconstructs fresh, role-specific contexts — sustain progress over long horizons more effectively than stateful agent frameworks. Key takeaways:

1. **Stateful search with stateless agents** is a viable and superior design for long-horizon AutoResearch
2. **Explicit coordination** (Advisor assignments) is both cheap (<0.6% of tokens) and effective
3. **Focused contexts** prevent the compounding degradation seen in long-running agent sessions
4. **Evaluation horizons matter**: short budgets systematically misjudge research systems and their components

**Future directions** include:
- Splitting the Advisor role across multiple stateless Advisors (possible because no conversation merging is needed)
- Using SLA trials as RL training episodes
- Extending to tasks with vague or slow-to-evaluate objectives

**Limitations**: Most configurations use single runs due to cost (hundreds to thousands of dollars per run); only Anthropic kernel optimization and checkpoint ablations are replicated; ablations change what agents see, not whether they keep conversations across invocations.

---

_Markdown view of https://picx.dev/p/707feb, served by PicX — AI-generated visual whiteboard summaries of research papers._
