Summary of "Similar Models Learn Differently: Final Window Pretraining Shapes Post-Training Beyond SFT"

Summary (Overview)

  • Core finding: Two model checkpoints that perform nearly identically after supervised fine-tuning (SFT) can respond very differently to subsequent preference optimization (DPO or RL), depending on the final window of pretraining data they were exposed to before instruction tuning.
  • Key methodology: Six branches fork from a shared 49B-token OLMo-2-1B checkpoint, differing only in a 500M-token final pretraining window (web text, filtered web text, normative discourse, safety text, math, or synthetic education). All branches then receive identical SFT and post-training.
  • Central result: The safety-text branch loses significantly less refusal behavior under DPO and GRPO post-training compared to the web-text branch, despite not starting with higher refusal after SFT — a "protection" effect of +8.2 points overall.
  • Selectivity: The effect is content-specific (only safety transformation text provides consistent cross-benchmark protection), requires the safety data to appear last in pretraining, and fades to near zero when the window shrinks to a few basis points of prior training.
  • Generalization: The divergence reproduces on a second model family (Pythia-1B) and under both DPO and GRPO (RL) updates, demonstrating that post-training is path-dependent on late pretraining in a way that standard post-SFT benchmarks cannot reveal.

Introduction and Theoretical Foundation

The paper addresses a fundamental question in LLM development: Are checkpoints that behave identically after SFT truly interchangeable for the next alignment stage?

The authors build on several strands of prior work:

  1. Training path matters: Models sharing a training trajectory converge to the same region of the loss landscape (Frankle et al. 2020; Neyshabur, Sedghi, and Zhang 2020), and plasticity — the capacity to keep changing under further training — is shaped by how a network was trained (Lyle et al. 2023; Dohare et al. 2024).

  2. Small pretraining data fractions matter: Target data mixed into pretraining at 1–2% rates changes downstream fine-tuning outcomes (Baek et al. 2026; Feng et al. 2026).

  3. Alignment-oriented text exists in pretraining corpora: Constitutions, preference feedback, and value surveys describe intended assistant behavior in natural language (Bai et al. 2022b; Kirk et al. 2024; Köpf et al. 2023), making it plausible that late pretraining on such material changes the state entering post-training.

The paper's hypothesis is that the final window of pretraining — the last data seen before instruction tuning — sets an "imprint" that determines how a model responds to post-training, even when post-SFT behavior appears identical. The key conceptual contribution is introducing refusal erosion as a stage-aware metric that separates how much refusal a branch starts with from how much of it the same update removes:

Eb(c)=Rb(θcS)−Rb(θcPT)(5)E_b(c) = R_b(\theta_c^S) - R_b(\theta_c^{PT})\tag{5}

This metric distinguishes between a branch ending with higher refusal because it started higher versus because it lost less during post-training.

Methodology

Experimental Design

Forked training setup: Starting from a shared partially-pretrained OLMo-2-1B checkpoint at ~49B tokens, the authors create six branches differing only in a 500M-token final continued pretraining (CPT) window:

θc=CPT(θ0,Dc,T)(1)\theta_c = \mathrm{CPT}(\theta_0, D_c, T)\tag{1}

The six corpora (Table 1) are designed to separate content effects:

  • CwebC_{web} (FineWeb-Edu): generic web text baseline
  • CdclmC_{dclm} (DCLM-baseline): filtered web text control
  • CnormC_{norm} (Normative): model governance prose, preference feedback, value surveys
  • CsafetyC_{safety} (Safety): harmful scenarios, unsafe drafts, safety critique
  • CmathC_{math} (OpenWebMath): math web text
  • CsynthC_{synth} (Cosmopedia): synthetic educational text

Post-training pipeline (identical across all branches):

  • SFT: 100k Tulu-style instruction examples including a fixed safety/refusal component
  • DPO: 60k UltraFeedback preference pairs, using the standard objective:
LDPO(θ)=−E(x,y+,y−)log⁡σ(β[log⁡πθ(y+∣x)πref(y+∣x)−log⁡πθ(y−∣x)πref(y−∣x)])(3)\mathcal{L}_{\mathrm{DPO}}(\theta) = -\mathbb{E}_{(x,y^+,y^-)} \log \sigma\left(\beta\left[\log \frac{\pi_\theta(y^+|x)}{\pi_{\mathrm{ref}}(y^+|x)} - \log \frac{\pi_\theta(y^-|x)}{\pi_{\mathrm{ref}}(y^-|x)}\right]\right)\tag{3}
  • GRPO: verifiable reward on GCD (greatest common divisor) task for the two headline branches

Order control: To isolate position effects, two branches train on the same DCLM and safety content in opposite orders (safety first vs. safety last), holding total tokens fixed.

Refusal Measurement

Refusal is classified via lexical string matching (ten refusal patterns within first 300 characters), validated against the WildGuard classifier with 97–99% agreement. Protection relative to the web baseline is defined as:

Pb(c)=Eb(Cweb)−Eb(c)(6)P_b(c) = E_b(C_{\mathrm{web}}) - E_b(c)\tag{6}

Empirical Validation / Results

1. Matched After SFT, Divergent Under DPO

After SFT, all branches are statistically indistinguishable on refusal, capability (MMLU, HellaSwag, ARC-Challenge, WinoGrande), and instruction following (IFEval) — within about one point, with CsafetyC_{safety} never ahead. Yet under identical DPO, the branches diverge dramatically:

BenchmarkEb(Cweb)E_b(C_{web})Eb(Csafety)E_b(C_{safety})Pb(Csafety)P_b(C_{safety})
AdvBench11.7 ± 0.38.4 ± 0.23.3 ± 0.5
BeaverTails (decontaminated)33.0 ± 1.527.8 ± 0.65.2 ± 1.2
XSTest-unsafe51.9 ± 1.235.7 ± 0.616.2 ± 1.2
Overall32.2 ± 0.824.0 ± 0.18.2 ± 0.9

Table 2: DPO refusal erosion. CsafetyC_{safety} loses less refusal than CwebC_{web} despite not starting higher — a difference in post-training susceptibility, not initial refusal strength.

2. Content Selectivity

The effect is not generic: DCLM and math are not protective; normative discourse is inconsistent (positive on some benchmarks, negative on others); synthetic education shows partial protection (strong on XSTest, weak on AdvBench). Only safety transformation text provides consistent cross-benchmark protection. Notably, the normative branch contains more explicit refusal language than safety but doesn't reproduce the effect, arguing against simple template mimicry.

3. Position Matters: Safety Must Come Last

The order control shows safety-last yields erosion of 2.6/25.2/34.8 on AdvBench/BeaverTails/XSTest-unsafe versus 12.7/44.2/50.7 for safety-first — a large gap that rules out an exposure-only explanation.

4. Generalization to RL (GRPO)

Under GRPO with a verifiable GCD reward, both branches learn the task equally (59.4% vs 58.7% exact accuracy), yet CwebC_{web} loses 16.9 ± 2.8 points of refusal while CsafetyC_{safety} loses only 6.1 ± 3.3 — a protection of 10.8 ± 1.0 points. A second task (Countdown arithmetic) reproduces this.

5. Robustness and Boundaries

  • Survives: lower DPO learning rate (2×10⁻⁷), alternative preference data (Chatbot Arena), Pythia-1B model family, decontaminated evaluation prompts
  • Fades with relative dose: protection drops from 9.1 points at 1.02% relative dose (49B fork) to 3.0 at 0.099% (504B) to ~0.5 at 0.0125% (4T checkpoint). A fixed-checkpoint control confirms ~0.1–1% relative dose is needed.

6. Trade-offs

CsafetyC_{safety} over-refuses benign prompts (OR-Bench-Hard) and carries a ~0.9 point capability cost already present at the SFT checkpoint — the effect is a broad shift in the refusal prior, not a calibrated improvement in harm refusal.

Theoretical and Practical Implications

Theoretical significance: This work demonstrates that post-training is path-dependent in a way that standard evaluation protocols cannot detect. The training trajectory leaves an imprint that affects plasticity for alignment, supporting and extending findings on how training history shapes network plasticity (Lyle et al. 2023; Dohare et al. 2024). The selectivity of the effect (safety text, not normative discourse or math) suggests the mechanism is task-proximal prior knowledge that interacts with later optimization, rather than generic exposure effects.

Practical implications:

  • Model evaluation: Post-SFT benchmarks (instruction following, refusal, capability) are insufficient to determine whether a checkpoint is ready for preference optimization. What a model was trained on last should be reported alongside its scores.
  • Alignment robustness: A small final pretraining window (~1% of training tokens) can meaningfully change how much SFT-installed safety behavior survives post-training — relevant for developers who acquire and align pre-trained checkpoints.
  • Data curation: The finding that safety-relevant text must appear last in pretraining to confer protection has concrete implications for pretraining data ordering.

Conclusion

The paper establishes that a small final window of pretraining shapes how a model responds to post-training beyond what SFT reveals. Safety text in the final window does not make a model refuse more after SFT; rather, it determines how much of that refusal survives the next update. The effect requires the safety data to come last, is selective to content, fades at smaller relative doses, and generalizes across models and algorithms (DPO and GRPO).

The central takeaway is that checkpoints should not be evaluated by post-SFT behavior alone — the training path leaves an imprint invisible to standard benchmarks yet decisive for how alignment proceeds. Future work could explore the mechanistic basis of this imprint, whether other SFT-installed behaviors show similar path dependence, and how to design pretraining curricula that deliberately shape post-training plasticity.

Related papers