Full text not available for this paper

Summary

  • Interior coverage optima: Every domain exhibits a non-monotonic (inverted-U) relationship between mid-training token coverage and downstream accuracy, with fitted peaks between 9.9% and 35.1% coverage—the moderate band (10–40%) is best for all five domains.
  • Alignment-resistant gaps: A fixed-budget compensatory SFT pass raises absolute accuracy in 116/120 cells (mean +4.32 pp) yet bridges only 0/240 pairwise gaps at a 5 pp threshold and 30/240 at a 10% ratio; an equal-budget uniform-SFT control behaves almost identically.
  • Transient negative transfer: Zero mid-training coverage collapses mid-training-only accuracy even where the base prior is high (Counterfactual 83.6% → 45.6%; Operation 60.4% → 33.2%), but the tested recipe partially repairs this.
  • Gap-preserving gains: A permutation null test (P < 0.001) shows the SFT gains are placed in a systematically gap-preserving way rather than randomly closing gaps.
  • Simplex confound: The five per-domain optima are mixture-level marginals, not coordinates of a jointly optimal mixture; the joint response surface's stationary point reads as a saddle in every domain.

Introduction and Theoretical Foundation

The paper addresses a critical gap in the multi-stage training paradigm (pre-training → mid-training → SFT → RL): per-domain data coverage at mid-training is typically set by data availability rather than principled design. The central question is whether a later alignment pass can undo coverage-associated differences.

The theoretical foundation rests on prior work showing:

  • Pretraining exposure determines RL generalization (Zhang et al., 2025)
  • RL amplifies pretrained behaviors (Zhao et al., 2025)
  • SFT primarily teaches format while RL generalizes from existing capabilities (Chu et al., 2025)
  • Data-mixture methods (DoReMi, RegMix) optimize domain proportions for aggregate perplexity

None of these systematically varies per-domain mid-training coverage while holding subsequent stages fixed—this is the gap the paper addresses.

The study uses KOR-Bench, a structured benchmark with five semantically rule-disjoint logical-reasoning domains (ciphers, custom mathematical operations, formal logic, constraint puzzles, counterfactual reasoning), which reduces semantic-transfer confounding. The primary model is Qwen3-8B-Base with a 4B replication.

Methodology

Training Stages

The pipeline consists of:

  1. Mid-training: Standard continued causal language modeling on corpus DmidD_{mid}, yielding checkpoint θmid\theta_{mid}
  2. Supervised Fine-Tuning (SFT): Optimizes conditional next-token likelihood over problem-response pairs, starting from θmid\theta_{mid}
  3. RL (GSPO): Group-based Sequence-level Policy Optimization with binary verifier rewards (Ri=1R_i = 1 if correct, 0 otherwise)

Experimental Design

  • 30 allocations spanning the five-domain simplex: 24 sweep configurations plus 6 withheld from the fit, at 5 seeds each
  • Fixed total token budget (≈1.5B tokens, two epochs) with a fixed external component (34.8% from ProofWriter rule family)
  • Coverage quantification: Measured by token share, not sample count
  • Data construction: Newly synthesized instances from benchmark rule definitions (not reused benchmark items), with strict split isolation

Key Analyses

  • Pairwise gap-closure analysis: Compensatory SFT reweights data toward coverage-deficient domains: rd∝max⁡(0,60%−Md)r_d \propto \max(0, 60\% - M_d)
  • Permutation null test: Reallocates observed gains at random to test whether gain placement preserves gaps
  • Out-of-sample validation: Six held-out allocations validate curve shapes (not peak locations)
  • Simplex-aware joint response surface: Maps allocations to isometric log-ratio coordinates and fits per-domain quadratic surfaces

Empirical Validation / Results

Finding 1: Interior Coverage Optima (Non-Monotonic Effects)

The moderate band (10–40%) is best for all five domains. Fitted split-Gaussian curves at 8B place peaks:

DomainFitted Peak
Cipher9.9%
Operation15.9%
Logic15.0%
Counterfactual26.0%
Puzzle35.1%

A calibrated permutation test for quadratic interiority gives P ≈ 0.010. Six withheld allocations reproduce curve shapes out-of-sample (residuals 1.0–3.0 pp), with Operation the weakest match (3.0 pp).

Finding 2: Persistent Coverage Gaps (Alignment-Resistant)

Table: Compensatory vs. Uniform SFT Gap Closure

MetricCompensatory SFTUniform SFT
Mean gain+4.32 pp+4.20 pp
Cells with gain116/120120/120
Pairs bridged (5 pp)0/2400/240
Pairs bridged (10% ratio)30/24032/240

The permutation null reallocating the same gains at random would bridge 13.8±3.313.8 \pm 3.3 and 77.9±8.577.9 \pm 8.5 pairs (P < 0.001), showing the gains are placed in a systematically gap-preserving way.

Sharpening trade-off: Raising the concentration exponent closes 0/60 → 12/60 pairs at the 5 pp metric, but mean gain falls +4.34 → +2.26 pp—closure and average accuracy are in tension under a fixed budget.

Full Pipeline Comparison (Table 1, selected rows)

ConfigOverallCipherOper.LogicCounterf.Puzzle
Base40.32 ± 2.76.8 ± 2.560.4 ± 2.947.2 ± 1.983.6 ± 3.03.6 ± 2.0
SFT+RL65.28 ± 1.968.0 ± 2.388.8 ± 3.059.2 ± 2.287.6 ± 3.022.8 ± 1.5
Expt. 1 (Imbalanced) Mid+SFT+RL65.92 ± 1.570.4 ± 2.192.8 ± 3.056.8 ± 2.088.8 ± 3.520.8 ± 2.0
Expt. 2 (Balanced) Mid+SFT+RL66.08 ± 2.372.0 ± 2.385.6 ± 3.260.8 ± 2.184.4 ± 3.327.6 ± 1.5
Expt. 3 (θ*) Mid+SFT+RL69.64 ± 2.663.0 ± 2.388.6 ± 2.575.9 ± 2.294.0 ± 3.326.7 ± 1.3

The exploratory θ* allocation attains the largest full-pipeline gain (+4.36 pp vs. +0.80/+0.64 pp for balanced/imbalanced) but is marginal under an uncorrected Welch test (t ≈ 2.29, p ≈ 0.052; t ≈ 2.77, p ≈ 0.030) and is selected from the same sweep.

Transient Negative Transfer

  • Operation: −27.2 pp vs. Base at mid-only → +5.2 pp over SFT-only baseline at mid+SFT (clear reversal)
  • Counterfactual: −38.0 pp at mid-only → still +0.8 pp below at mid+SFT → +1.2 pp by RL (not statistically resolved)
  • The FineWeb-Edu-only control shows the collapse is co-mingled with generic distributional drift: mid-training on unrelated data drops Counterfactual from 83.6% to 29.2% (−54.4 pp), a larger drop than the zero-coverage case.

External Benchmarks (Limited Consistency Checks)

BenchmarkSpearman ρ with KOR average
ProofWriter+0.34
ZebraLogic Grid+0.53
ZebraLogic MC−0.12
CounterBench+0.67

These are descriptive (n = 9), not evidence of a transfer mechanism.

Theoretical and Practical Implications

Theoretical Implications

  1. Coverage effects are non-monotonic: The inverted-U pattern challenges the assumption that more domain-specific data is always better; the moderate band (10–40%) is consistently optimal.

  2. Alignment has limited corrective power: The structural negative result (gaps survive alignment) suggests that mid-training coverage choices create constraints that later stages cannot easily undo—extending Zhang et al. (2025) to the mid-training setting.

  3. Simplex confound is fundamental: The five per-domain optima are marginal statements, not a jointly achievable mixture. The joint response surface's stationary point reads as a saddle in every domain (provisional classification given the surface's parameter count).

Practical Implications

  • Coverage allocation may warrant explicit auditing in mid-training design, rather than being left to data availability.
  • The θ gain (+4.36 pp) is marginal* under an uncorrected Welch test and remains a model-selected candidate from the same sweep—allocation rules should optimise predicted accuracy directly rather than relative-gain objectives.
  • The tested recipe's limited dynamic range (arithmetic cap of ≈7 pp by construction) means the compensation formula family has inherent constraints.

Conclusion

Three findings hold in this setting:

  1. Coverage-induced domain gaps survive finite-budget post-training: Compensatory and uniform SFT both leave gaps essentially intact (0/240 bridged at 5 pp), generalizing to held-out allocations, with gains placed in a gap-preserving way (P < 0.001).

  2. Zero coverage collapses mid-training-only accuracy even with high priors, but the tested recipe reverses the signs (partially), leaving small residual deficits on Logic and Puzzle; the collapse is co-mingled with generic distributional drift.

  3. Every domain has an interior coverage optimum: Moderate band (10–40%) yields highest mean accuracy for all five domains, with fitted peaks between 9.9% and 35.1%; these are per-domain marginals, not a jointly achievable mixture.

Future directions include:

  • Separating the simplex confound via a joint response-surface model with more allocations
  • Testing whether other alignment policy families or larger budgets can close gaps
  • Investigating the capability-envelope hypothesis (representations, gradient conflict, parameter overlap)
  • Developing principled allocation rules that optimise predicted accuracy directly

The paper's central caveat: "coverage allocation may warrant explicit auditing, but the θ* gain is marginal under an uncorrected Welch test and remains a model-selected candidate from the same sweep." The findings support a hypothesis to test rather than a rule to follow.

Related papers