Full text not available for this paper
Summary (Overview)
- T1 is a 122B-parameter Mixture-of-Experts (MoE) terminal agent trained purely via reinforcement learning (PPO) on executed outcomes in real Linux cloud sandboxes, operating for up to 300+ tool-call turns per task.
- The training pipeline achieves 64.0% resolved on Terminal-Bench 2.1, a 28.5% relative gain over the SFT checkpoint (49.4%), surpassing GPT-5.4 (54.8%), DeepSeek-V4-Flash (56.9%), and approaching Claude Opus 4.7 (66.1%) with only 10B active parameters.
- Two novel stabilization mechanisms—TITO (Token-In-Token-Out) and R³ (Rollout Routing Replay)—cut the training-to-inference log-probability gap from 0.021 to 0.013 with zero token drift in the loss region.
- A dense verification reward based on absolute passing assertion counts (rather than binary outcomes) enables learning from partial progress; the first binary-reward campaign never exceeded its supervised baseline.
- The training corpus (T1-15k) is fully out-of-distribution from evaluation benchmarks, ensuring gains reflect genuine capability transfer rather than benchmark overfitting.
Introduction and Theoretical Foundation
The paper addresses the shift in agentic AI from single-turn tasks to autonomous long-horizon execution, where models must issue actions with persistent consequences in stateful environments, verified by execution rather than preference. The Linux terminal is positioned as the "sharpest and most unforgiving test" because it binds abstract planning to irreversible side effects.
The core theoretical challenge is training-inference consistency for sparse MoE models:
- Expert weights account for 116.0B of 121.4B parameters; each token engages 8 of 256 experts per layer through a discrete router.
- Minor numeric differences between inference and training stacks can flip expert selections, causing gradients to reach different parameters than those that generated the behavior.
- Multi-turn harnesses perturb the token sequence at every turn boundary.
The paper formalizes two independent fidelity conditions for stable RL:
Token fidelity: — the token identifiers used in training must match those sampled during rollout.
Routing fidelity: — the expert routing masks must match.
The policy is written as where indexes the sub-network of 95.5% of parameters held by experts, with for Qwen3.5-122B-A10B.
Methodology
Training Framework Overview
The pipeline runs asynchronous PPO with:
- Inference replicas (SGLang) serving the behavior policy under oversampling
- Training backend (Megatron) updating critic then actor
- Daytona sandboxes executing multi-turn commands for reward collection
Dataset Construction (T1-15k)
Three pools are used:
| Pool | Size | Verifier Type | Purpose |
|---|---|---|---|
| TMax-15k | 14,601 | Binary only | Critic warm-up |
| RST-38k | 37,484 | Per-assertion | Unfiltered dense reward |
| T1-15k | 15,000 | Per-assertion | Production training |
Tasks are selected via an 8-dimension weighted audit using DeepSeek-V4-Pro:
| Dimension | Weight | Facet |
|---|---|---|
| Instruction/Verifier Alignment | 0.20 | Instruction |
| Verifier Fairness | 0.15 | Verifier |
| Solution Correctness | 0.15 | Solution |
| Instruction Clarity | 0.10 | Instruction |
| Instruction Self-Contained | 0.10 | Instruction |
| Verifier Coverage | 0.10 | Verifier |
| Solution Reasonableness | 0.10 | Solution |
| Task Training Value | 0.10 | Task Value |
Hard gates reject tasks with hidden requirements, test leakage, solution shortcuts, or weak verifiers.
TITO: Token-In-Token-Out
TITO ensures the trainer consumes exactly the token identifiers the sampler emitted. The stream is with loss mask and log-probabilities satisfying:
Boundary repair cases (from strongest to weakest):
- Strict: , context =
- Normalized: s.t. ,
- Retokenized: , by offset mapping
- Split: new chunk under same trial
The drift rate inside the loss region is measured at exactly 0.0000%:
R³: Rollout Routing Replay
R³ records the routing mask inference selected and replays it in the training forward pass. The training pass keeps normalization on live logits but takes selection from the recorded mask:
The cost is negligible: (1.5 KiB per position), keeping rollout overhead below 3%.
Dense Verification Reward
The reward is the absolute passing count on a fixed global scale:
where is the number of passing assertions and is chosen near the 90th percentile of the assertion-count distribution (median = 4, max ≈ 35). This is explicitly not a pass ratio—passing 10/20 assertions on a hard task gives while passing 2/4 on an easy task gives .
Critic Warm-Up and PPO Configuration
The critic is warm-started via one epoch of training on TMax-15k before any policy step. PPO uses:
- (GAE)
- (actor clip), (critic clip)
- Critic learning rate: (30× the actor's )
- Load balancing coefficient = 0, KL penalties disabled
- Adam with , weight decay 0.1
The advantage estimator and value loss are:
Empirical Validation / Results
Main Results on Terminal-Bench 2.1
| Model | Size (total-active) | Terminal-Bench 2.1 | LHTB |
|---|---|---|---|
| Claude Opus 4.7 | / | 66.1 | – |
| T1 | 122B-A10B | 64.0 | 27.9 |
| Claude Opus 4.6 | / | 63.8 | – |
| Muse Spark | / | 62.2 | – |
| RST-38k RL | 122B-A10B | 59.9 | 25.4 |
| Hy3-Preview | 295B-A21B | 58.0 | – |
| DeepSeek V4 Flash | 295B-A21B | 56.9 | – |
| GPT-5.4 | / | 54.8 | 27.2 |
| RST-SFT Model | 122B-A10B | 49.4 | 23.6 |
| Qwen3.5-122B-A10B (base) | 122B-A10B | 43.8 | 18.9 |
Training Dynamics
- Step 10 already reaches 56.2% (+6.8 points over SFT initialization)
- Steps 20–60 hold a band of 55.1–57.3%
- Step 70 reaches 61.8%, step 110 peaks at 64.0%
- RL contributes 14.6 percentage points vs. 5.6 for SFT — 72.3% of total gain from RL
Transfer to Longer Horizons
- Long-Horizon Terminal Bench: T1 scores 27.9, matching Gemini-3.1-Pro, exceeding GPT-5.4 (27.2) and GLM-5.1 (26.7)
- Terminal-Bench Hard: T1 resolves 38.0%, exceeding DeepSeek-V4-Pro (36.0%) and beating SFT by 9.7 points
Domain-Specific Gains
- Debugging: T1 = 100.0% vs. GPT-5.6 Sol = 80.0%
- System Administration: T1 = 88.9% vs. GPT-5.6 Sol = 55.6%
- All checkpoints reach 100% on Easy tasks; Medium: 78/58/56%; Hard: 33/30/20%
Stability Measurements
- Training-inference log-probability gap: 0.013 with TITO+R³ vs. 0.021 without
- Critic explained variance: 0.71–0.86 with warm-up vs. −33.6 cold start
- Average turns per trajectory: 10.4 → 20.9 (bounded, not runaway)
- Total sequence length: 11.5k → 17.8k tokens
Theoretical and Practical Implications
Theoretical Contributions
-
Two-axis decomposition of training-inference mismatch: The paper formalizes that sparse MoE RL requires both token fidelity and routing fidelity as independent conditions, each addressable by separate mechanisms.
-
Dense reward from execution: Demonstrates that absolute passing counts on a fixed global scale provide learnable signal where binary rewards fail, with the critic serving as the only baseline (no group statistics exist with single-sample-per-task).
-
Capacity model for co-resident actor-critic pairs: Proposition 7.1 shows expert optimizer state is plan-invariant—the dominant memory term depends only on device count , not on how experts are distributed:
Practical Implications
-
Recipe portability: The capacity model admits plans analytically before initialization, making the recipe portable across device budgets spanning more than 2×.
-
Infrastructure insights: Context parallelism for recurrent operators (sequence-sharding) enables 84k/128k context walls; liveness under hour-scale steps requires bounded deadlines on control-plane operations but not bulk transfers.
-
Failure mode documentation: The paper candidly reports that critic calibration alone doesn't guarantee policy improvement, context management determines learnability, and rollout throughput changes the training distribution.
-
GRPO vs. PPO: Group-relative baselines degenerate for long-horizon tasks because groups of all-fail or all-succeed trajectories contribute no gradient, and group closure waits on the slowest member.
Conclusion
T1 demonstrates that pure reinforcement learning on executed outcomes can train a 122B MoE terminal agent to frontier-level performance with only 10B active parameters. The key takeaways:
-
TITO + R³ are the critical stabilization mechanisms for sparse MoE RL, reducing training-inference mismatch from 0.021 to 0.013 with exactly aligned zero token drift.
-
Dense verification rewards (absolute passing counts on a fixed scale) unlock learning from partial progress, which binary rewards cannot provide.
-
Critic warm-up is essential: cold-start critics begin at EV = −33.6 and spend half the campaign recovering, while warm-started critics settle at 0.71–0.86 immediately.
-
Out-of-distribution training ensures gains reflect genuine capability transfer, not benchmark overfitting.
Future Directions
- Eliminating residual re-tokenization: 2.6% of tokens in the audited run are trained on sampled identifiers where inference used re-tokenized ones
- Single-axis reward ablation: Isolating binary vs. dense reward contributions with matched pool, initialization, and routing replay
- Verifier integrity enforcement: In-sandbox tamper detection (read-only test mounts, checksummed verifiers)
- Partial-rollout continuation for the cancelled oversampling tail
- Broader data distribution to cover machine learning/data-science/scientific-computing categories
- Fully-asynchronous training with staleness-robust combinations (GSPO+SAT+R³, soft masking, TIS)
- Critic value pretraining on offline trajectories and classification-based value losses (HL-Gauss)
Related papers
- Growing an Agent/Prover Interface: Evolutionary Tool Design for Cost-Efficient Theorem Proving in Rocq and Lean
Evolutionary interface design lets a frontier model mutate agent-prover tools, yielding an MCP server that boosts theorem-proving accuracy by up to 15 points while cutting costs and time by over half.
- False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
CrossFit, a cross-fitted feedback method that scores proposals with auxiliary solvers trained on complementary source folds, eliminates co-cheating in self-evolving search agents, boosting downstream performance by 8.8 points.
- Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
Identical runs of the same AI coding agent vary more than differences between agent-model pairings, so best-of-three with compliance checking beats single-run benchmarking.