The Value of Repeated Tokens Collapses to a Single Coordinate
Highlights of This Issue
The clearest signal this issue is that the value of repeated tokens is unified onto a single coordinate. What Is a Repeated Token Worth anchors repeated tokens to two explicit baselines (same-data single epoch and equal-compute fresh data), finds that excess loss collapses to a single variable z=(R-1)N_total/U, and gives a critical epoch number R_c≈2.7(D/N)^0.24—the number of epochs at which the value of repetition drops to half of fresh data rises with per-parameter training budget and barely changes with model scale (What Is a Repeated Token Worth?). This unified explanation resolves the seemingly conflicting scaling trends under three designs—fixed corpus, fixed U/N, and fixed TPP—and provides a compute-optimal allocation path under a fixed unique dataset: larger models and more epochs grow together until loss stops improving. It advances the repetition-damage research from Issue 1's Internal Data Repetition and Issue 2's Prescriptive Scaling Laws into an actionable geometric framework, directly changing allocation decisions under data-constrained settings.
The post-training ceiling criterion saw advances in two directions this issue. On one hand, Final Window Pretraining proves that the final pretraining window leaves an invisible imprint between checkpoints that match in behavior after SFT, determining the extent to which subsequent DPO/RL erodes the behavior installed by SFT—the safety-text branch does not show higher refusal rates after SFT, yet loses less refusal under the same post-training updates (Final Window Pretraining). This builds on Issue 5's Good Pretraining, Bad SFT, where base loss optimality does not equal post-training optimality, forming a progression: even if matching after SFT, pretraining path dependence can still differentiate post-training endpoints. On the other hand, What pass@k Cannot Measure points out that pass@k itself has a structural blind spot as a diversity proxy—it depends only on the probability of correct samples, cannot perceive reorganization of the output distribution, and the lack of a control relative to the starting checkpoint can cause GRPO's pass@1 advantage to be misread as a capability gain (What pass@k Cannot Measure).
Data selection mechanisms saw multiple new levers this issue. Fisher-Guided Submodular reverses Fisher information from an EWC-style regularizer into a CPT data selector, providing evidence for parameter-space curvature as a selection criterion with 1B selected tokens outperforming 10B replay at 10x token efficiency (Fisher-Guided Submodular Data Selection). LESSER reduces the cost of extracting gradient features for post-training data selection by an order of magnitude, covering SFT/RL/distillation settings (LESSER). On synthetic data, It's All Training proposes a fully synthetic, single-stage (collapsing pre/mid/post-training) recipe where models exhibit instruction-following and reasoning without SFT/RL (It's All Training), while How Far Can Synthetic Data Take Thai OCR advances the boundary of synthetic data benefits to which attributes of synthetic rendering truly transfer, and provides counterintuitive evidence that training granularity (page vs crop) reverses the relative advantage of in-domain/out-of-domain (How Far Can Synthetic Data Take Thai OCR?).
Open Questions
- At what scale does the single coordinate z and critical epoch number R_c from What Is a Repeated Token Worth hold? Under its 2B upper bound and base loss criterion, can the conclusion that repetition value rises with per-parameter budget and barely changes with scale be reproduced at ≥1B with post-training ceiling as the criterion?
- Does Final Window Pretraining's finding that matching after SFT can still diverge depend on its single behavioral probe of refusal rate? If switched to capability-type (math/code) final windows, does the path-dependence imprint still exist, and can it be incorporated into data allocation decisions in mid-training/annealing stages?
- Whether Fisher geometry as a CPT data selection signal outperforms loss/perplexity selectors depends on its reliance on warmup checkpoints and diagonal approximation—at larger scales and longer CPT, does the mechanistic explanation of Fisher diagonal drift still hold?
- Does pass@k's structural blind spot imply that the post-training ceiling criterion itself needs a diversity dimension? If pass@k cannot distinguish output distribution reorganization, what protocol should evaluation of data recipe quality shift to?
Papers in this issue
Excess loss from data repetition follows a single power law in (R-1)N/U, unifying conflicting size trends and enabling compute-optimal epoch allocation.
Editor's noteFirst to anchor the value and cost of repeated tokens to two explicit baselines (same-data single epoch and equal-compute fresh data), finding that excess loss collapses to a single coordinate z=(R-1)N_total/U, and giving a critical epoch number R_c≈2.7(D/N)^0.24—the number of epochs at which repetition value halves rises with per-parameter budget and barely changes with model scale. Relative to Issue 1's Internal Data Repetition and Issue 2's Prescriptive Scaling Laws, the increment is reading repetition value from measured single-epoch curves rather than fitted laws, and unified explanation of seemingly conflicting scaling trends under fixed corpus, fixed U/N, and fixed TPP designs, directly giving a compute-optimal allocation path under a fixed unique dataset. Limitation: the criterion is still base loss and code is not released, but as a geometric unified framework for repetition-damage research it is worth reading.
Identical post-SFT checkpoints diverge dramatically under preference optimization, with safety text in the final pretraining window preserving refusal behavior by up to 8.2 points.
Editor's noteFirst to systematically prove that final-window pretraining leaves an invisible imprint between checkpoints that match in behavior after SFT, determining the extent to which subsequent DPO/RL erodes the behavior installed by SFT: the safety-text branch does not show higher refusal rates after SFT, yet loses less refusal under the same post-training updates. Relative to Issue 5's Good Pretraining, Bad SFT, which treats base loss optimality not equaling post-training optimality as a diagnosis, the increment is explicitly advancing pretraining path dependence into an actionable criterion of matching after SFT can still diverge, with content selectivity, order dependence, and relative dose boundaries. At 1B real scale, six-branch controlled ablations, and both DPO and GRPO updates, it is a strong design constraint for mid-training/annealing data selection.
Fisher-guided gradient decomposition with submodular selection achieves 10x token efficiency over replay in continual pretraining, dominating Pareto frontiers on forgetting and adaptation.
Editor's noteReverses Fisher information from an EWC-style regularizer into a CPT data selector: decomposes each candidate gradient into an anchor component along high-Fisher directions and a frontier component along low-Fisher directions, and aggregates via a log-det submodular objective in a streaming single pass. Relative to Issue 8's ReScraper (extraction quality axis) and Issue 6's MiST (mid-training corpus design), the increment is making parameter-space curvature an explicit criterion for CPT data selection, with evidence of 1B selected tokens outperforming 10B replay at 10x token efficiency. At 1.1B/8B real scale, with controlled comparisons against DSIR/DoReMi/EWC, it directly targets mid-training data selection and forgetting control.
SYNTH's fully synthetic pre-training corpus collapses pre-, mid-, and post-training into one stage, achieving 10-140x token efficiency and state-of-the-art factual precision across all model sizes.
Editor's noteFirst to propose a fully synthetic, single-stage (collapsing pre/mid/post-training) pretraining recipe: from 58k Wikipedia seeds, generates an 80B-token corpus via constrained grammar and dual auxiliary model back-translation, where models exhibit instruction-following and reasoning without SFT/RL. Relative to Issue 6's QVAC Genesis III (student failure signals) and Issue 1's Generating Pretraining Tokens (model-aware generation), the increment is explicitly using constrained priors to control query diversity and seed grounding to control memorization, with evidence at 13B MoE scale and FActScore factual precision. Public corpus and model, directly addressing the boundary of synthetic data benefits and single-stage training feasibility.
Synthetic document reconstruction trains a competitive Thai OCR model without real Thai labels, with typeface diversity and 2D layout driving transfer more than page context.
Editor's noteFirst to systematically decouple multiple factors in synthetic data transfer on Thai OCR (source domain, non-text context, font diversity, 2D layout, handwriting glyph source), and finds that training granularity (page-level vs crop-level) reverses the relative advantage of in-domain vs out-of-domain reconstruction—in-domain reconstruction approaches real supervision under page-level (1.82% vs 1.31% CER) but degrades sharply under crop-level (15.59% vs 5.52%). Relative to Issue 6's WeVisDoc and Issue 1's PureDocBench, the increment advances the document extraction quality axis to controlled factor decomposition of synthetic data transfer, with direct evidence of training granularity as a moderating variable. Limitation: relatively small scale (0.9B–2B) and CER as criterion.
LESSER replaces expensive full-parameter gradients with output-layer gradients computed from forward passes, cutting feature-extraction FLOPs by up to 9.7× while matching full-gradient data selection performance within 1.3 points.
Editor's noteFirst to systematically prove that output-layer gradients can serve as an efficient substitute for full-parameter gradient features in three types of post-training data selection—SFT, RL, and distillation teacher selection—and can be extracted with only one forward pass. Relative to Issue 7's Data-DPO target-model-conditional selection, the increment is making the computational cost of gradient features itself the object of study, giving 9.7x/3.0x FLOP reduction and a batch-gradient alignment mechanism explanation—even if single-sample rankings disagree, the selected batch's gradient still aligns with the full-gradient method. Covering SFT/RL/distillation settings, it directly addresses post-training ceiling and data selection cost concerns.
pass@k cannot detect diversity loss: GRPO and RFT move entropy and answer diversity in opposite directions with zero seed overlap, yet pass@8 and pass@32 show no consistent winner.
Editor's noteFirst to systematically prove that pass@k has a structural blind spot in post-training evaluation: it depends only on the probability of correct samples and cannot perceive diversity changes in the output distribution. GRPO vs RFT comparison shows three diversity metrics (token entropy, answer entropy, unique answer count) moving in opposite directions, while pass@8/32 cannot distinguish them; and relative to the starting checkpoint, GRPO's pass@1 advantage is only a smaller loss, not a capability gain. Relative to Issue 5's Good Pretraining, Bad SFT, which focuses on the disconnect between base loss and post-training ceiling, the increment targets pass@k itself as a failed diversity proxy, with methodological value for studies relying on pass@k to judge data recipe quality.
The Kontoyiannis entropy rate estimator, a model-free text filter, outperforms logprob-based filtering by boosting unique trigrams 42% and cutting repetition 19% during iterative fine-tuning.
Editor's noteFirst to propose using the Kontoyiannis entropy rate estimator (a nonparametric statistic based on match length) as a fully model-free text filter to mitigate model collapse in iterative fine-tuning, and proves it outperforms surplexity filtering baselines that require model logprobs. Relative to Issue 8's How Much Is an AI Token Worth? on scaling laws for wild AI text, this work focuses on data filtering in iterative fine-tuning scenarios and does not rely on any model access—zero cost, no GPU needed, providing another detection and intervention perspective for synthetic data collapse. Limitation: criterion is text diversity rather than post-training ceiling.







