Full text not available for this paper
Summary
- Interior coverage optima: Every domain exhibits a non-monotonic (inverted-U) relationship between mid-training token coverage and downstream accuracy, with fitted peaks between 9.9% and 35.1% coverage—the moderate band (10–40%) is best for all five domains.
- Alignment-resistant gaps: A fixed-budget compensatory SFT pass raises absolute accuracy in 116/120 cells (mean +4.32 pp) yet bridges only 0/240 pairwise gaps at a 5 pp threshold and 30/240 at a 10% ratio; an equal-budget uniform-SFT control behaves almost identically.
- Transient negative transfer: Zero mid-training coverage collapses mid-training-only accuracy even where the base prior is high (Counterfactual 83.6% → 45.6%; Operation 60.4% → 33.2%), but the tested recipe partially repairs this.
- Gap-preserving gains: A permutation null test (P < 0.001) shows the SFT gains are placed in a systematically gap-preserving way rather than randomly closing gaps.
- Simplex confound: The five per-domain optima are mixture-level marginals, not coordinates of a jointly optimal mixture; the joint response surface's stationary point reads as a saddle in every domain.
Introduction and Theoretical Foundation
The paper addresses a critical gap in the multi-stage training paradigm (pre-training → mid-training → SFT → RL): per-domain data coverage at mid-training is typically set by data availability rather than principled design. The central question is whether a later alignment pass can undo coverage-associated differences.
The theoretical foundation rests on prior work showing:
- Pretraining exposure determines RL generalization (Zhang et al., 2025)
- RL amplifies pretrained behaviors (Zhao et al., 2025)
- SFT primarily teaches format while RL generalizes from existing capabilities (Chu et al., 2025)
- Data-mixture methods (DoReMi, RegMix) optimize domain proportions for aggregate perplexity
None of these systematically varies per-domain mid-training coverage while holding subsequent stages fixed—this is the gap the paper addresses.
The study uses KOR-Bench, a structured benchmark with five semantically rule-disjoint logical-reasoning domains (ciphers, custom mathematical operations, formal logic, constraint puzzles, counterfactual reasoning), which reduces semantic-transfer confounding. The primary model is Qwen3-8B-Base with a 4B replication.
Methodology
Training Stages
The pipeline consists of:
- Mid-training: Standard continued causal language modeling on corpus , yielding checkpoint
- Supervised Fine-Tuning (SFT): Optimizes conditional next-token likelihood over problem-response pairs, starting from
- RL (GSPO): Group-based Sequence-level Policy Optimization with binary verifier rewards ( if correct, 0 otherwise)
Experimental Design
- 30 allocations spanning the five-domain simplex: 24 sweep configurations plus 6 withheld from the fit, at 5 seeds each
- Fixed total token budget (≈1.5B tokens, two epochs) with a fixed external component (34.8% from ProofWriter rule family)
- Coverage quantification: Measured by token share, not sample count
- Data construction: Newly synthesized instances from benchmark rule definitions (not reused benchmark items), with strict split isolation
Key Analyses
- Pairwise gap-closure analysis: Compensatory SFT reweights data toward coverage-deficient domains:
- Permutation null test: Reallocates observed gains at random to test whether gain placement preserves gaps
- Out-of-sample validation: Six held-out allocations validate curve shapes (not peak locations)
- Simplex-aware joint response surface: Maps allocations to isometric log-ratio coordinates and fits per-domain quadratic surfaces
Empirical Validation / Results
Finding 1: Interior Coverage Optima (Non-Monotonic Effects)
The moderate band (10–40%) is best for all five domains. Fitted split-Gaussian curves at 8B place peaks:
| Domain | Fitted Peak |
|---|---|
| Cipher | 9.9% |
| Operation | 15.9% |
| Logic | 15.0% |
| Counterfactual | 26.0% |
| Puzzle | 35.1% |
A calibrated permutation test for quadratic interiority gives P ≈ 0.010. Six withheld allocations reproduce curve shapes out-of-sample (residuals 1.0–3.0 pp), with Operation the weakest match (3.0 pp).
Finding 2: Persistent Coverage Gaps (Alignment-Resistant)
Table: Compensatory vs. Uniform SFT Gap Closure
| Metric | Compensatory SFT | Uniform SFT |
|---|---|---|
| Mean gain | +4.32 pp | +4.20 pp |
| Cells with gain | 116/120 | 120/120 |
| Pairs bridged (5 pp) | 0/240 | 0/240 |
| Pairs bridged (10% ratio) | 30/240 | 32/240 |
The permutation null reallocating the same gains at random would bridge and pairs (P < 0.001), showing the gains are placed in a systematically gap-preserving way.
Sharpening trade-off: Raising the concentration exponent closes 0/60 → 12/60 pairs at the 5 pp metric, but mean gain falls +4.34 → +2.26 pp—closure and average accuracy are in tension under a fixed budget.
Full Pipeline Comparison (Table 1, selected rows)
| Config | Overall | Cipher | Oper. | Logic | Counterf. | Puzzle |
|---|---|---|---|---|---|---|
| Base | 40.32 ± 2.7 | 6.8 ± 2.5 | 60.4 ± 2.9 | 47.2 ± 1.9 | 83.6 ± 3.0 | 3.6 ± 2.0 |
| SFT+RL | 65.28 ± 1.9 | 68.0 ± 2.3 | 88.8 ± 3.0 | 59.2 ± 2.2 | 87.6 ± 3.0 | 22.8 ± 1.5 |
| Expt. 1 (Imbalanced) Mid+SFT+RL | 65.92 ± 1.5 | 70.4 ± 2.1 | 92.8 ± 3.0 | 56.8 ± 2.0 | 88.8 ± 3.5 | 20.8 ± 2.0 |
| Expt. 2 (Balanced) Mid+SFT+RL | 66.08 ± 2.3 | 72.0 ± 2.3 | 85.6 ± 3.2 | 60.8 ± 2.1 | 84.4 ± 3.3 | 27.6 ± 1.5 |
| Expt. 3 (θ*) Mid+SFT+RL | 69.64 ± 2.6 | 63.0 ± 2.3 | 88.6 ± 2.5 | 75.9 ± 2.2 | 94.0 ± 3.3 | 26.7 ± 1.3 |
The exploratory θ* allocation attains the largest full-pipeline gain (+4.36 pp vs. +0.80/+0.64 pp for balanced/imbalanced) but is marginal under an uncorrected Welch test (t ≈ 2.29, p ≈ 0.052; t ≈ 2.77, p ≈ 0.030) and is selected from the same sweep.
Transient Negative Transfer
- Operation: −27.2 pp vs. Base at mid-only → +5.2 pp over SFT-only baseline at mid+SFT (clear reversal)
- Counterfactual: −38.0 pp at mid-only → still +0.8 pp below at mid+SFT → +1.2 pp by RL (not statistically resolved)
- The FineWeb-Edu-only control shows the collapse is co-mingled with generic distributional drift: mid-training on unrelated data drops Counterfactual from 83.6% to 29.2% (−54.4 pp), a larger drop than the zero-coverage case.
External Benchmarks (Limited Consistency Checks)
| Benchmark | Spearman ρ with KOR average |
|---|---|
| ProofWriter | +0.34 |
| ZebraLogic Grid | +0.53 |
| ZebraLogic MC | −0.12 |
| CounterBench | +0.67 |
These are descriptive (n = 9), not evidence of a transfer mechanism.
Theoretical and Practical Implications
Theoretical Implications
-
Coverage effects are non-monotonic: The inverted-U pattern challenges the assumption that more domain-specific data is always better; the moderate band (10–40%) is consistently optimal.
-
Alignment has limited corrective power: The structural negative result (gaps survive alignment) suggests that mid-training coverage choices create constraints that later stages cannot easily undo—extending Zhang et al. (2025) to the mid-training setting.
-
Simplex confound is fundamental: The five per-domain optima are marginal statements, not a jointly achievable mixture. The joint response surface's stationary point reads as a saddle in every domain (provisional classification given the surface's parameter count).
Practical Implications
- Coverage allocation may warrant explicit auditing in mid-training design, rather than being left to data availability.
- The θ gain (+4.36 pp) is marginal* under an uncorrected Welch test and remains a model-selected candidate from the same sweep—allocation rules should optimise predicted accuracy directly rather than relative-gain objectives.
- The tested recipe's limited dynamic range (arithmetic cap of ≈7 pp by construction) means the compensation formula family has inherent constraints.
Conclusion
Three findings hold in this setting:
-
Coverage-induced domain gaps survive finite-budget post-training: Compensatory and uniform SFT both leave gaps essentially intact (0/240 bridged at 5 pp), generalizing to held-out allocations, with gains placed in a gap-preserving way (P < 0.001).
-
Zero coverage collapses mid-training-only accuracy even with high priors, but the tested recipe reverses the signs (partially), leaving small residual deficits on Logic and Puzzle; the collapse is co-mingled with generic distributional drift.
-
Every domain has an interior coverage optimum: Moderate band (10–40%) yields highest mean accuracy for all five domains, with fitted peaks between 9.9% and 35.1%; these are per-domain marginals, not a jointly achievable mixture.
Future directions include:
- Separating the simplex confound via a joint response-surface model with more allocations
- Testing whether other alignment policy families or larger budgets can close gaps
- Investigating the capability-envelope hypothesis (representations, gradient conflict, parameter overlap)
- Developing principled allocation rules that optimise predicted accuracy directly
The paper's central caveat: "coverage allocation may warrant explicit auditing, but the θ* gain is marginal under an uncorrected Welch test and remains a model-selected candidate from the same sweep." The findings support a hypothesis to test rather than a rule to follow.
Related papers
- How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Under matched conditions, a minimal-harness coding agent matches or outperforms state-of-the-art MLE harnesses, with performance driven by the LLM backbone and execution environment, not scaffolding.
- Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
SuffixReplay enables fine-grained prefix caching in hybrid LLMs by reconstructing linear-attention states from sparse anchors, cutting storage by 2x and TTFT by up to 70%.
- How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
Encoder-free multimodal LLMs match encoder-based performance at ~10^22 FLOPs, shifting compute-optimal allocation toward larger decoders and enabling viable encoder-free pretraining.