Full text not available for this paper

Summary (Overview)

  • T1 is a 122B-parameter Mixture-of-Experts (MoE) terminal agent trained purely via reinforcement learning (PPO) on executed outcomes in real Linux cloud sandboxes, operating for up to 300+ tool-call turns per task.
  • The training pipeline achieves 64.0% resolved on Terminal-Bench 2.1, a 28.5% relative gain over the SFT checkpoint (49.4%), surpassing GPT-5.4 (54.8%), DeepSeek-V4-Flash (56.9%), and approaching Claude Opus 4.7 (66.1%) with only 10B active parameters.
  • Two novel stabilization mechanisms—TITO (Token-In-Token-Out) and R³ (Rollout Routing Replay)—cut the training-to-inference log-probability gap from 0.021 to 0.013 with zero token drift in the loss region.
  • A dense verification reward based on absolute passing assertion counts (rather than binary outcomes) enables learning from partial progress; the first binary-reward campaign never exceeded its supervised baseline.
  • The training corpus (T1-15k) is fully out-of-distribution from evaluation benchmarks, ensuring gains reflect genuine capability transfer rather than benchmark overfitting.

Introduction and Theoretical Foundation

The paper addresses the shift in agentic AI from single-turn tasks to autonomous long-horizon execution, where models must issue actions with persistent consequences in stateful environments, verified by execution rather than preference. The Linux terminal is positioned as the "sharpest and most unforgiving test" because it binds abstract planning to irreversible side effects.

The core theoretical challenge is training-inference consistency for sparse MoE models:

  • Expert weights account for 116.0B of 121.4B parameters; each token engages 8 of 256 experts per layer through a discrete router.
  • Minor numeric differences between inference and training stacks can flip expert selections, causing gradients to reach different parameters than those that generated the behavior.
  • Multi-turn harnesses perturb the token sequence at every turn boundary.

The paper formalizes two independent fidelity conditions for stable RL:

Token fidelity: Tjt=TjrT^t_j = T^r_j — the token identifiers used in training must match those sampled during rollout.

Routing fidelity: Ijt,ℓ=Ijr,ℓ ∀ℓ∈[L]I^{t,\ell}_j = I^{r,\ell}_j \ \forall \ell \in [L] — the expert routing masks must match.

The policy is written as πθ(⋅∣T<j,Ij)\pi_\theta(\cdot | T_{<j}, I_j) where Ij=(Ij1,…,IjL)I_j = (I^1_j, \ldots, I^L_j) indexes the sub-network of 95.5% of parameters held by experts, with (L,k,E)=(48,8,256)(L, k, E) = (48, 8, 256) for Qwen3.5-122B-A10B.


Methodology

Training Framework Overview

The pipeline runs asynchronous PPO with:

  • Inference replicas (SGLang) serving the behavior policy under oversampling
  • Training backend (Megatron) updating critic then actor
  • Daytona sandboxes executing multi-turn commands for reward collection

Dataset Construction (T1-15k)

Three pools are used:

PoolSizeVerifier TypePurpose
TMax-15k14,601Binary onlyCritic warm-up
RST-38k37,484Per-assertionUnfiltered dense reward
T1-15k15,000Per-assertionProduction training

Tasks are selected via an 8-dimension weighted audit using DeepSeek-V4-Pro:

DimensionWeightFacet
Instruction/Verifier Alignment0.20Instruction
Verifier Fairness0.15Verifier
Solution Correctness0.15Solution
Instruction Clarity0.10Instruction
Instruction Self-Contained0.10Instruction
Verifier Coverage0.10Verifier
Solution Reasonableness0.10Solution
Task Training Value0.10Task Value

Hard gates reject tasks with hidden requirements, test leakage, solution shortcuts, or weak verifiers.

TITO: Token-In-Token-Out

TITO ensures the trainer consumes exactly the token identifiers the sampler emitted. The stream is T=(T1,…,TN)T = (T_1, \ldots, T_N) with loss mask mm and log-probabilities qq satisfying:

mj=1  ⟺  Tj was sampled by πt−1r,qj=mj⋅log⁡πt−1r(Tj∣T<j,Ijr)m_j = 1 \iff T_j \text{ was sampled by } \pi^r_{t-1}, \quad q_j = m_j \cdot \log \pi^r_{t-1}\left(T_j | T_{<j}, I^r_j\right)

Boundary repair cases (from strongest to weakest):

  1. Strict: Ti⪯pi+1T_i \preceq p_{i+1}, context = pi+1⊖Tip_{i+1} \ominus T_i
  2. Normalized: min⁡(s,u)(s+u)\min_{(s,u)}(s+u) s.t. drops(pi)∥dropu(ai)⪯pi+1\text{drop}_s(p_i) \| \text{drop}_u(a_i) \preceq p_{i+1}, (s,u)∈[0,96]×[0,16](s,u) \in [0,96] \times [0,16]
  3. Retokenized: dec(ai)=dec(a~i)\text{dec}(a_i) = \text{dec}(\tilde{a}_i), a~i=pi+1[b:e]\tilde{a}_i = p_{i+1}[b:e] by offset mapping
  4. Split: new chunk under same trial

The drift rate inside the loss region is measured at exactly 0.0000%:

∣{j∈D∪P:mj=1}∣∣{j:mj=1}∣=0.0000%\frac{|\{j \in D \cup P : m_j = 1\}|}{|\{j : m_j = 1\}|} = 0.0000\%

R³: Rollout Routing Replay

R³ records the routing mask inference selected and replays it in the training forward pass. The training pass keeps normalization on live logits but takes selection from the recorded mask:

gj,eℓ=exp⁡(sjℓ(θt))e∑e′∈Ijr,ℓexp⁡(sjℓ(θt))e′ for e∈Ijr,ℓ,yjℓ=∑e∈Ijr,ℓgj,eℓfe(hjℓ)g^\ell_{j,e} = \frac{\exp(s^\ell_j(\theta_t))_e}{\sum_{e' \in I^{r,\ell}_j} \exp(s^\ell_j(\theta_t))_{e'}} \ \text{for } e \in I^{r,\ell}_j, \quad y^\ell_j = \sum_{e \in I^{r,\ell}_j} g^\ell_{j,e} f_e(h^\ell_j)

The cost is negligible: btok=Lk⋅4 B=1536 B/tokenb_{\text{tok}} = Lk \cdot 4 \text{ B} = 1536 \text{ B/token} (1.5 KiB per position), keeping rollout overhead below 3%.

Dense Verification Reward

The reward is the absolute passing count on a fixed global scale:

r=PS,S=20r = \frac{P}{S}, \quad S = 20

where PP is the number of passing assertions and S=20S = 20 is chosen near the 90th percentile of the assertion-count distribution (median = 4, max ≈ 35). This is explicitly not a pass ratio—passing 10/20 assertions on a hard task gives 10/20=0.510/20 = 0.5 while passing 2/4 on an easy task gives 2/20=0.12/20 = 0.1.

Critic Warm-Up and PPO Configuration

The critic is warm-started via one epoch of training on TMax-15k before any policy step. PPO uses:

  • γ=λ=1\gamma = \lambda = 1 (GAE)
  • ϵ=0.2\epsilon = 0.2 (actor clip), ϵv=0.2\epsilon_v = 0.2 (critic clip)
  • Critic learning rate: 1.5×10−51.5 \times 10^{-5} (30× the actor's 1.0×10−61.0 \times 10^{-6})
  • Load balancing coefficient = 0, KL penalties disabled
  • Adam with β=(0.9,0.98)\beta = (0.9, 0.98), weight decay 0.1

The advantage estimator and value loss are:

A^j=∑n≥0(γλ)nδj+n,where δj=r^j+γVold(sj+1)−Vold(sj)\hat{A}_j = \sum_{n \geq 0}(\gamma\lambda)^n \delta_{j+n}, \quad \text{where } \delta_j = \hat{r}_j + \gamma V^{\text{old}}(s_{j+1}) - V^{\text{old}}(s_j) LV(ϕ)=Ej[max⁡((Vϕ(sj)−R^j)2,(Vclip(sj)−R^j)2)]L_V(\phi) = \mathbb{E}_j\left[\max\left((V_\phi(s_j) - \hat{R}_j)^2, (V^{\text{clip}}(s_j) - \hat{R}_j)^2\right)\right] Lπ(θ)=−Ej[min⁡(rj(θ)A^j,clip(rj(θ),1−ϵ,1+ϵ)A^j)]L_\pi(\theta) = -\mathbb{E}_j\left[\min\left(r_j(\theta)\hat{A}_j, \text{clip}(r_j(\theta), 1-\epsilon, 1+\epsilon)\hat{A}_j\right)\right]

Empirical Validation / Results

Main Results on Terminal-Bench 2.1

ModelSize (total-active)Terminal-Bench 2.1LHTB
Claude Opus 4.7/66.1–
T1122B-A10B64.027.9
Claude Opus 4.6/63.8–
Muse Spark/62.2–
RST-38k RL122B-A10B59.925.4
Hy3-Preview295B-A21B58.0–
DeepSeek V4 Flash295B-A21B56.9–
GPT-5.4/54.827.2
RST-SFT Model122B-A10B49.423.6
Qwen3.5-122B-A10B (base)122B-A10B43.818.9

Training Dynamics

  • Step 10 already reaches 56.2% (+6.8 points over SFT initialization)
  • Steps 20–60 hold a band of 55.1–57.3%
  • Step 70 reaches 61.8%, step 110 peaks at 64.0%
  • RL contributes 14.6 percentage points vs. 5.6 for SFT — 72.3% of total gain from RL

Transfer to Longer Horizons

  • Long-Horizon Terminal Bench: T1 scores 27.9, matching Gemini-3.1-Pro, exceeding GPT-5.4 (27.2) and GLM-5.1 (26.7)
  • Terminal-Bench Hard: T1 resolves 38.0%, exceeding DeepSeek-V4-Pro (36.0%) and beating SFT by 9.7 points

Domain-Specific Gains

  • Debugging: T1 = 100.0% vs. GPT-5.6 Sol = 80.0%
  • System Administration: T1 = 88.9% vs. GPT-5.6 Sol = 55.6%
  • All checkpoints reach 100% on Easy tasks; Medium: 78/58/56%; Hard: 33/30/20%

Stability Measurements

  • Training-inference log-probability gap: 0.013 with TITO+R³ vs. 0.021 without
  • Critic explained variance: 0.71–0.86 with warm-up vs. −33.6 cold start
  • Average turns per trajectory: 10.4 → 20.9 (bounded, not runaway)
  • Total sequence length: 11.5k → 17.8k tokens

Theoretical and Practical Implications

Theoretical Contributions

  1. Two-axis decomposition of training-inference mismatch: The paper formalizes that sparse MoE RL requires both token fidelity and routing fidelity as independent conditions, each addressable by separate mechanisms.

  2. Dense reward from execution: Demonstrates that absolute passing counts on a fixed global scale provide learnable signal where binary rewards fail, with the critic serving as the only baseline (no group statistics exist with single-sample-per-task).

  3. Capacity model for co-resident actor-critic pairs: Proposition 7.1 shows expert optimizer state is plan-invariant—the dominant memory term depends only on device count GG, not on how experts are distributed:

Per-device expert optimizer state∝1η×τκδη×η=G\text{Per-device expert optimizer state} \propto \frac{1}{\eta} \times \frac{\tau\kappa\delta}{\eta} \times \eta = G

Practical Implications

  1. Recipe portability: The capacity model admits plans analytically before initialization, making the recipe portable across device budgets spanning more than 2×.

  2. Infrastructure insights: Context parallelism for recurrent operators (sequence-sharding) enables 84k/128k context walls; liveness under hour-scale steps requires bounded deadlines on control-plane operations but not bulk transfers.

  3. Failure mode documentation: The paper candidly reports that critic calibration alone doesn't guarantee policy improvement, context management determines learnability, and rollout throughput changes the training distribution.

  4. GRPO vs. PPO: Group-relative baselines degenerate for long-horizon tasks because groups of all-fail or all-succeed trajectories contribute no gradient, and group closure waits on the slowest member.


Conclusion

T1 demonstrates that pure reinforcement learning on executed outcomes can train a 122B MoE terminal agent to frontier-level performance with only 10B active parameters. The key takeaways:

  1. TITO + R³ are the critical stabilization mechanisms for sparse MoE RL, reducing training-inference mismatch from 0.021 to 0.013 with exactly aligned zero token drift.

  2. Dense verification rewards (absolute passing counts on a fixed scale) unlock learning from partial progress, which binary rewards cannot provide.

  3. Critic warm-up is essential: cold-start critics begin at EV = −33.6 and spend half the campaign recovering, while warm-started critics settle at 0.71–0.86 immediately.

  4. Out-of-distribution training ensures gains reflect genuine capability transfer, not benchmark overfitting.

Future Directions

  • Eliminating residual re-tokenization: 2.6% of tokens in the audited run are trained on sampled identifiers where inference used re-tokenized ones
  • Single-axis reward ablation: Isolating binary vs. dense reward contributions with matched pool, initialization, and routing replay
  • Verifier integrity enforcement: In-sandbox tamper detection (read-only test mounts, checksummed verifiers)
  • Partial-rollout continuation for the cancelled oversampling tail
  • Broader data distribution to cover machine learning/data-science/scientific-computing categories
  • Fully-asynchronous training with staleness-robust combinations (GSPO+SAT+R³, soft masking, TIS)
  • Critic value pretraining on offline trajectories and classification-based value losses (HL-Gauss)

Related papers