# T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

> Pure RL on executed outcomes trains a 122B-parameter MoE terminal agent to 64.0% on Terminal-Bench 2.1, surpassing GPT-5.4 with only 10B active parameters.

- **Source:** [arXiv](https://arxiv.org/abs/2609.11042)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/gARPI0
- **Whiteboard:** https://picx.dev/p/gARPI0/image

## Summary

## Summary (Overview)

- **T1** is a 122B-parameter Mixture-of-Experts (MoE) terminal agent trained purely via reinforcement learning (PPO) on executed outcomes in real Linux cloud sandboxes, operating for up to 300+ tool-call turns per task.
- The training pipeline achieves **64.0% resolved on Terminal-Bench 2.1**, a **28.5% relative gain** over the SFT checkpoint (49.4%), surpassing GPT-5.4 (54.8%), DeepSeek-V4-Flash (56.9%), and approaching Claude Opus 4.7 (66.1%) with only 10B active parameters.
- Two novel stabilization mechanisms—**TITO** (Token-In-Token-Out) and **R³** (Rollout Routing Replay)—cut the training-to-inference log-probability gap from 0.021 to 0.013 with zero token drift in the loss region.
- A **dense verification reward** based on absolute passing assertion counts (rather than binary outcomes) enables learning from partial progress; the first binary-reward campaign never exceeded its supervised baseline.
- The training corpus (T1-15k) is fully out-of-distribution from evaluation benchmarks, ensuring gains reflect genuine capability transfer rather than benchmark overfitting.

---

## Introduction and Theoretical Foundation

The paper addresses the shift in agentic AI from single-turn tasks to **autonomous long-horizon execution**, where models must issue actions with persistent consequences in stateful environments, verified by execution rather than preference. The Linux terminal is positioned as the "sharpest and most unforgiving test" because it binds abstract planning to irreversible side effects.

The core theoretical challenge is **training-inference consistency for sparse MoE models**:
- Expert weights account for 116.0B of 121.4B parameters; each token engages 8 of 256 experts per layer through a discrete router.
- Minor numeric differences between inference and training stacks can flip expert selections, causing gradients to reach different parameters than those that generated the behavior.
- Multi-turn harnesses perturb the token sequence at every turn boundary.

The paper formalizes two independent fidelity conditions for stable RL:

**Token fidelity:** $T^t_j = T^r_j$ — the token identifiers used in training must match those sampled during rollout.

**Routing fidelity:** $I^{t,\ell}_j = I^{r,\ell}_j \ \forall \ell \in [L]$ — the expert routing masks must match.

The policy is written as $\pi_\theta(\cdot | T_{<j}, I_j)$ where $I_j = (I^1_j, \ldots, I^L_j)$ indexes the sub-network of 95.5% of parameters held by experts, with $(L, k, E) = (48, 8, 256)$ for Qwen3.5-122B-A10B.

---

## Methodology

### Training Framework Overview

The pipeline runs **asynchronous PPO** with:
- **Inference replicas** (SGLang) serving the behavior policy under oversampling
- **Training backend** (Megatron) updating critic then actor
- **Daytona sandboxes** executing multi-turn commands for reward collection

### Dataset Construction (T1-15k)

Three pools are used:
| Pool | Size | Verifier Type | Purpose |
|------|------|---------------|---------|
| TMax-15k | 14,601 | Binary only | Critic warm-up |
| RST-38k | 37,484 | Per-assertion | Unfiltered dense reward |
| T1-15k | 15,000 | Per-assertion | Production training |

Tasks are selected via an **8-dimension weighted audit** using DeepSeek-V4-Pro:

| Dimension | Weight | Facet |
|-----------|--------|-------|
| Instruction/Verifier Alignment | 0.20 | Instruction |
| Verifier Fairness | 0.15 | Verifier |
| Solution Correctness | 0.15 | Solution |
| Instruction Clarity | 0.10 | Instruction |
| Instruction Self-Contained | 0.10 | Instruction |
| Verifier Coverage | 0.10 | Verifier |
| Solution Reasonableness | 0.10 | Solution |
| Task Training Value | 0.10 | Task Value |

Hard gates reject tasks with hidden requirements, test leakage, solution shortcuts, or weak verifiers.

### TITO: Token-In-Token-Out

TITO ensures the trainer consumes **exactly the token identifiers the sampler emitted**. The stream is $T = (T_1, \ldots, T_N)$ with loss mask $m$ and log-probabilities $q$ satisfying:

$$m_j = 1 \iff T_j \text{ was sampled by } \pi^r_{t-1}, \quad q_j = m_j \cdot \log \pi^r_{t-1}\left(T_j | T_{<j}, I^r_j\right)$$

Boundary repair cases (from strongest to weakest):
1. **Strict:** $T_i \preceq p_{i+1}$, context = $p_{i+1} \ominus T_i$
2. **Normalized:** $\min_{(s,u)}(s+u)$ s.t. $\text{drop}_s(p_i) \| \text{drop}_u(a_i) \preceq p_{i+1}$, $(s,u) \in [0,96] \times [0,16]$
3. **Retokenized:** $\text{dec}(a_i) = \text{dec}(\tilde{a}_i)$, $\tilde{a}_i = p_{i+1}[b:e]$ by offset mapping
4. **Split:** new chunk under same trial

The drift rate inside the loss region is measured at exactly **0.0000%**:

$$\frac{|\{j \in D \cup P : m_j = 1\}|}{|\{j : m_j = 1\}|} = 0.0000\%$$

### R³: Rollout Routing Replay

R³ records the routing mask inference selected and replays it in the training forward pass. The training pass keeps normalization on live logits but takes selection from the recorded mask:

$$g^\ell_{j,e} = \frac{\exp(s^\ell_j(\theta_t))_e}{\sum_{e' \in I^{r,\ell}_j} \exp(s^\ell_j(\theta_t))_{e'}} \ \text{for } e \in I^{r,\ell}_j, \quad y^\ell_j = \sum_{e \in I^{r,\ell}_j} g^\ell_{j,e} f_e(h^\ell_j)$$

The cost is negligible: $b_{\text{tok}} = Lk \cdot 4 \text{ B} = 1536 \text{ B/token}$ (1.5 KiB per position), keeping rollout overhead below 3%.

### Dense Verification Reward

The reward is the **absolute passing count on a fixed global scale**:

$$r = \frac{P}{S}, \quad S = 20$$

where $P$ is the number of passing assertions and $S = 20$ is chosen near the 90th percentile of the assertion-count distribution (median = 4, max ≈ 35). This is explicitly **not** a pass ratio—passing 10/20 assertions on a hard task gives $10/20 = 0.5$ while passing 2/4 on an easy task gives $2/20 = 0.1$.

### Critic Warm-Up and PPO Configuration

The critic is warm-started via one epoch of training on TMax-15k before any policy step. PPO uses:
- $\gamma = \lambda = 1$ (GAE)
- $\epsilon = 0.2$ (actor clip), $\epsilon_v = 0.2$ (critic clip)
- Critic learning rate: $1.5 \times 10^{-5}$ (30× the actor's $1.0 \times 10^{-6}$)
- Load balancing coefficient = 0, KL penalties disabled
- Adam with $\beta = (0.9, 0.98)$, weight decay 0.1

The advantage estimator and value loss are:

$$\hat{A}_j = \sum_{n \geq 0}(\gamma\lambda)^n \delta_{j+n}, \quad \text{where } \delta_j = \hat{r}_j + \gamma V^{\text{old}}(s_{j+1}) - V^{\text{old}}(s_j)$$

$$L_V(\phi) = \mathbb{E}_j\left[\max\left((V_\phi(s_j) - \hat{R}_j)^2, (V^{\text{clip}}(s_j) - \hat{R}_j)^2\right)\right]$$

$$L_\pi(\theta) = -\mathbb{E}_j\left[\min\left(r_j(\theta)\hat{A}_j, \text{clip}(r_j(\theta), 1-\epsilon, 1+\epsilon)\hat{A}_j\right)\right]$$

---

## Empirical Validation / Results

### Main Results on Terminal-Bench 2.1

| Model | Size (total-active) | Terminal-Bench 2.1 | LHTB |
|-------|---------------------|--------------------|------|
| Claude Opus 4.7 | / | 66.1 | – |
| **T1** | **122B-A10B** | **64.0** | **27.9** |
| Claude Opus 4.6 | / | 63.8 | – |
| Muse Spark | / | 62.2 | – |
| RST-38k RL | 122B-A10B | 59.9 | 25.4 |
| Hy3-Preview | 295B-A21B | 58.0 | – |
| DeepSeek V4 Flash | 295B-A21B | 56.9 | – |
| GPT-5.4 | / | 54.8 | 27.2 |
| RST-SFT Model | 122B-A10B | 49.4 | 23.6 |
| Qwen3.5-122B-A10B (base) | 122B-A10B | 43.8 | 18.9 |

### Training Dynamics

- **Step 10** already reaches 56.2% (+6.8 points over SFT initialization)
- **Steps 20–60** hold a band of 55.1–57.3%
- **Step 70** reaches 61.8%, **step 110** peaks at 64.0%
- RL contributes 14.6 percentage points vs. 5.6 for SFT — **72.3% of total gain from RL**

### Transfer to Longer Horizons

- **Long-Horizon Terminal Bench:** T1 scores 27.9, matching Gemini-3.1-Pro, exceeding GPT-5.4 (27.2) and GLM-5.1 (26.7)
- **Terminal-Bench Hard:** T1 resolves 38.0%, exceeding DeepSeek-V4-Pro (36.0%) and beating SFT by 9.7 points

### Domain-Specific Gains

- **Debugging:** T1 = 100.0% vs. GPT-5.6 Sol = 80.0%
- **System Administration:** T1 = 88.9% vs. GPT-5.6 Sol = 55.6%
- All checkpoints reach 100% on Easy tasks; Medium: 78/58/56%; Hard: 33/30/20%

### Stability Measurements

- Training-inference log-probability gap: **0.013** with TITO+R³ vs. **0.021** without
- Critic explained variance: **0.71–0.86** with warm-up vs. **−33.6** cold start
- Average turns per trajectory: 10.4 → 20.9 (bounded, not runaway)
- Total sequence length: 11.5k → 17.8k tokens

---

## Theoretical and Practical Implications

### Theoretical Contributions

1. **Two-axis decomposition of training-inference mismatch:** The paper formalizes that sparse MoE RL requires both token fidelity and routing fidelity as independent conditions, each addressable by separate mechanisms.

2. **Dense reward from execution:** Demonstrates that absolute passing counts on a fixed global scale provide learnable signal where binary rewards fail, with the critic serving as the only baseline (no group statistics exist with single-sample-per-task).

3. **Capacity model for co-resident actor-critic pairs:** Proposition 7.1 shows expert optimizer state is **plan-invariant**—the dominant memory term depends only on device count $G$, not on how experts are distributed:

$$\text{Per-device expert optimizer state} \propto \frac{1}{\eta} \times \frac{\tau\kappa\delta}{\eta} \times \eta = G$$

### Practical Implications

1. **Recipe portability:** The capacity model admits plans analytically before initialization, making the recipe portable across device budgets spanning more than 2×.

2. **Infrastructure insights:** Context parallelism for recurrent operators (sequence-sharding) enables 84k/128k context walls; liveness under hour-scale steps requires bounded deadlines on control-plane operations but not bulk transfers.

3. **Failure mode documentation:** The paper candidly reports that critic calibration alone doesn't guarantee policy improvement, context management determines learnability, and rollout throughput changes the training distribution.

4. **GRPO vs. PPO:** Group-relative baselines degenerate for long-horizon tasks because groups of all-fail or all-succeed trajectories contribute no gradient, and group closure waits on the slowest member.

---

## Conclusion

T1 demonstrates that **pure reinforcement learning on executed outcomes** can train a 122B MoE terminal agent to frontier-level performance with only 10B active parameters. The key takeaways:

1. **TITO + R³** are the critical stabilization mechanisms for sparse MoE RL, reducing training-inference mismatch from 0.021 to 0.013 with exactly aligned zero token drift.

2. **Dense verification rewards** (absolute passing counts on a fixed scale) unlock learning from partial progress, which binary rewards cannot provide.

3. **Critic warm-up** is essential: cold-start critics begin at EV = −33.6 and spend half the campaign recovering, while warm-started critics settle at 0.71–0.86 immediately.

4. **Out-of-distribution training** ensures gains reflect genuine capability transfer, not benchmark overfitting.

### Future Directions

- **Eliminating residual re-tokenization:** 2.6% of tokens in the audited run are trained on sampled identifiers where inference used re-tokenized ones
- **Single-axis reward ablation:** Isolating binary vs. dense reward contributions with matched pool, initialization, and routing replay
- **Verifier integrity enforcement:** In-sandbox tamper detection (read-only test mounts, checksummed verifiers)
- **Partial-rollout continuation** for the cancelled oversampling tail
- **Broader data distribution** to cover machine learning/data-science/scientific-computing categories
- **Fully-asynchronous training** with staleness-robust combinations (GSPO+SAT+R³, soft masking, TIS)
- **Critic value pretraining** on offline trajectories and classification-based value losses (HL-Gauss)

---

_Markdown view of https://picx.dev/p/gARPI0, served by PicX — AI-generated visual whiteboard summaries of research papers._
