No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse
Author: Lewis Mitchell, Adelaide Data Science Centre, School of Mathematical Sciences, Adelaide University
Summary (Overview)
- Novel model-free collapse filter: The paper introduces the Kontoyiannis entropy rate estimator as a training-data filter against model collapse—the first completely model-free and reference-free filter proposed, requiring no log-probabilities, no external oracle, and no real human data.
- Superior performance over logprob-based filtering: In a six-generation QLoRA fine-tuning experiment on Llama-3.1-8B, -filtering yields +42% unique trigrams, +30% vocabulary, and −19% repetition (all ), while logprob-based filtering () shows no significant text-diversity benefit on any metric ().
- Validated as cross-domain entropy proxy: correlates strongly with logprob entropy (, ) across 4 domains, 2 temperatures, and 2 generator–scorer model pairs, with a domain-invariant slope.
- Effective collapse detector: Under fine-tuning collapse, within-topic Spearman concordance between and is () across 39 topics.
- Practical and efficient: Scoring a 1,500-word document takes under 10 ms with no GPU, making collapse-resistant curation feasible for closed-source APIs, federated training, and cross-organisation data sharing.
Introduction and Theoretical Foundation
Background and Motivation
Large language models trained on their own outputs face model collapse: a self-reinforcing degradation where output entropy declines across generations, diversity narrows, and rare linguistic patterns are progressively lost (Shumailov et al., 2024; Alemohammad et al., 2024). As synthetic content becomes an increasing fraction of web-scraped training corpora, collapse mitigation is shifting from theoretical concern to engineering requirement.
Existing Mitigation Categories
| Category | Approach | Requirement |
|---|---|---|
| (A) Model-logprob methods | Surplexity filtering (Gambetta et al., 2026); token resampling (Zhu et al., 2025); top-K sampling | Access to model's output distribution |
| (B) External-oracle methods | Verifier-based filtering (Feng et al., 2025; Yi et al., 2025) | Stronger external model or human annotator |
| (C) Real-data accumulation | Data accumulation (Gerstgrasser et al., 2024; Bertrand et al., 2024; Fu et al., 2025) | Continued access to human-authored data |
Each category imposes an extra burden. Surplexity filtering—the most established model-access-requiring baseline—requires a forward pass through the model being trained, is unavailable for closed-source APIs, computationally expensive at scale, and version-dependent. Notably, Guo et al. (2024) find that linguistic acceptability filtering can actually worsen collapse, underscoring that filter choice is consequential.
The Kontoyiannis Entropy Rate Estimator
Grounded in mathematical information theory, the paper proposes the Kontoyiannis entropy rate estimator (Kontoyiannis et al., 1998):
where is the length of the shortest prefix of that does not appear as a contiguous substring in . This converges almost surely to the entropy rate of any stationary ergodic process—no parametric assumptions required.
Key properties:
- Match-length statistics: Long matches signal repetition; short matches signal novelty
- LZ compression connection: Each is the elementary step of Lempel–Ziv parsing, making interpretable as a normalised LZ compression rate
- Consistency: Grounded in the Shannon–McMillan–Breiman theorem: almost surely
- No model, no API, no GPU required
Prior applications include quantifying information flow in social media text (Bagrow et al., 2019; Pond et al., 2020), but no prior work has applied as a collapse filter.
Methodology
Experimental Design
Three parallel 6-generation QLoRA fine-tuning chains on Llama-3.1-8B-Instruct:
- 80 documents generated per generation across 4 domains (encyclopedic, creative, scientific, conversational; 10 topics × 2 documents per topic per domain)
- 40 selected for training according to condition:
- Unfiltered: random selection (seed fixed per generation)
- -filter: top 40 by mean per-token Shannon entropy (logprob-based)
- -filter: top 40 by Kontoyiannis entropy rate (text-only)
- Domain-stratified selection (10 documents per domain)
- Generation-0 documents shared across conditions (drawn from base model)
Training configuration:
- 4-bit NF4 QLoRA (rank 16, α = 16)
- Learning rate , cosine schedule, 3 epochs per generation
- Effective batch size 8
- Fixed decoding temperature of 1.0
Tokenization: The ProcessEntropy package's tokenizer (NLTK TweetTokenizer pass, lower-cased, non-alphanumeric characters stripped). Real-time filtering uses faster whitespace tokenization (Spearman agreement).
Finite-sample bias control: computed on first 1,500 word tokens; shorter texts discarded.
Validation Experiments
- Cross-domain/cross-temperature validation: 200 documents from GPT-4o across 4 domains, 2 temperatures (), 25 topics per domain; 197 passed reliability filter (<5% unrecoverable logprobs)
- Cross-model validation: Gemini 1.5 Pro generated additional corpora; Qwen2.5-7B-Instruct used as independent scorer
- Collapse detection: Within-topic concordance across generations under both rephrasing and fine-tuning regimes
Empirical Validation / Results
1. as Proxy for Logprob Entropy
Cross-domain result: OLS regression of on with domain fixed effects:
- Likelihood-ratio test for domain-specific slopes fails to reject equality (): domain-invariant slope
- Domains differ in baseline entropy (creative/conversational > scientific) but scales with at the same rate in each
Cross-model result: Both GPT-4o/Qwen and Gemini/Qwen pairings reproduce the positive relationship, confirming measures a property of the text itself, not any particular model's distribution.
Scope limitation: On generation-0 data from Llama-3.1-8B-Instruct, correlation is near zero (, , ). The base model produces a wide bimodal distribution (SD = 2.17) driven by topic-dependent familiarity rather than document-level diversity—this is the root cause of -filtering's failure.
2. Collapse Detection Without Logprobs
| Regime | Within-generation concordance | Interpretation |
|---|---|---|
| Iterative rephrasing | Spearman | Rephrasing homogenises surface statistics |
| Fine-tuning collapse | () | Phrase-level repetition directly measured by |
Under rephrasing, and co-decline at indistinguishable rates (, ), but fine-tuning collapse changes surface statistics in ways directly measures.
3. Training Data Filter Results (Generation 6)
Table 1: Generation-6 outcomes under three training-data conditions (n = 80 documents per condition; p-values from Welch t-tests)
| Metric | Unfilt. | -filt. | -filt. | |||
|---|---|---|---|---|---|---|
| (bits/word) | 0.79 | 1.16 | 1.43 | <0.001 | 0.008 | <0.001 |
| Distinct-3 (%) | — | +7% | +42% | <0.001 | 0.23 | <0.001 |
| Vocabulary (%) | — | +4% | +30% | <0.001 | 0.31 | <0.001 |
| Rep-4 (%) | — | −3% | −19% | <0.001 | 0.27 | <0.001 |
Key findings:
- -filtering retains 1.43 bits/word vs. 0.79 unfiltered ()
- -filtering provides no text-diversity benefit ( on all diversity metrics). Mechanistic explanation: distribution compresses dramatically after one fine-tuning step (gen-0 SD = 2.17 → gen-1 SD = 0.13), collapsing the ranking signal
- Filter selection sets differ substantially: within-domain Jaccard similarity ≈ 0.35 (range 0.18–0.54), indicating complementary axes
4. Comparison with Simpler Text-Only Baselines
Static-pool comparison on generation-0 documents (exact comparison, all conditions share same 80 documents):
| Filter | Mean of selected docs | Jaccard vs. -filter |
|---|---|---|
| -filtering | 3.98 | — |
| TTR@500-filtering | 3.37 | 0.43 |
| Distinct-2-filtering | 3.32 | — |
| gzip-compression ratio | — | 0.36–0.38 |
| Repetition-rate filter | — | 0.36–0.38 |
As genuine six-generation training filters, -filtering significantly outperforms both gzip-filtering (Distinct-3 +17%, vocabulary +17%, Rep-4 −7%; all ) and repetition-rate filtering (no benefit; all ) on every metric (all , most ).
5. Downstream Benchmarks
- HellaSwag: Accuracy drops from 0.729 (gen-0) to ≈0.652–0.657 (gen-6) with no significant difference between conditions—-filtering preserves text diversity but not general reasoning ability
- MMLU (5-shot): 67.6–67.6% across conditions at gen-6 (all pairwise tests )
- GSM8K (8-shot): 70.7–71.2% across conditions at gen-6 (all )
- MAUVE scores: Near-ceiling and uniform (0.987–0.993), indicating embedding-space similarity preserved
6. Quality Assessment
LLM-judge pass (GPT-4o-mini, blind to condition) on all 240 generation-6 documents:
| Metric | -filt. | Unfiltered | -filt. |
|---|---|---|---|
| Coherence (1–5) | 2.35 | 2.20 () | 2.29 |
| Instruction-following (1–5) | 3.23 | 3.00 () | 3.15 |
-filtering's diversity advantage comes with a small quality benefit rather than a coherence cost.
7. Replication with Smaller Model
With Llama-3.2-3B-Instruct (same corpus and protocol):
- -filtering retains 2.91 bits/word vs. 2.12 unfiltered () and 2.41 for -filtering ()
- -filtering again non-significant on all text-diversity metrics
- -vs- comparison reaches significance on the metric ( vs. at 8B)
Theoretical and Practical Implications
Why -filtering outperforms -filtering
-
Direct targeting of collapse signature: directly penalises repeated phrase structure—the primary surface manifestation of fine-tuning collapse. Documents with recurring n-grams exhibit long match lengths and hence low .
-
signal destruction: measures token-level uncertainty in the scorer model—a signal destroyed after a single fine-tuning step. At generation-0, primarily reflects topic familiarity rather than document diversity.
-
Surplexity collapse: A literal reimplementation shows median absolute deviation falls from 4.40 (gen-0 pool) to 0.40–0.49 after one fine-tuning step—a 9–11-fold narrowing.
Simpson's Paradox and the Jaccard Gap
The two filters select different documents (Jaccard ≈ 0.35) despite corpus-level correlation () due to a form of Simpson's paradox:
- Between-domain correlation is strong (creative/conversational have both high and high )
- Within-domain residual correlation is near zero—where filtering operates due to domain stratification
The estimators measure genuinely different properties of intra-domain text variation.
Practical Implications
- Zero-cost filtering: Scoring a 1,500-word document takes under 10 ms; a pipeline producing candidate documents per round incurs negligible, trivially parallelisable overhead
- Applicable where logprob access is unavailable: closed-source API distillation, cross-organisation data sharing, federated training
- Complementary to deduplication: Min-hash near-deduplication removes between-document duplicates; -filtering ensures within-document structure hasn't collapsed. Combined pipeline (deduplicate first, then -filter) addresses both failure modes
- Multi-agent diversity: The Kontoyiannis cross-entropy rate could operationalise the blind-writing independence phase Chen et al. (2026) identify as the most effective structural intervention against diversity collapse
Conclusion
Main Takeaways
The Kontoyiannis entropy rate estimator —a model-free, match-length measure of sequential complexity—is an effective training-data filter against fine-tuning-driven model collapse:
- +42% unique trigrams, +30% vocabulary, −19% intra-document repetition vs. unfiltered training at generation 6
- Logprob-based filtering shows no significant diversity benefit ( on all metrics)
- Zero-cost, model-free, reference-free: no model access, no API calls, no GPU required
Key Limitations
- Scope: -filtering preserves text diversity but does not protect general reasoning ability (HellaSwag, MMLU, GSM8K all degrade uniformly)
- Specificity: Advantage is specific to fine-tuning collapse, not iterative rephrasing (where within-generation rank concordance is ≈ 0)
- Lexical focus: Diversity gains are lexical (MAUVE scores near-ceiling and uniform across conditions)
- Single model family: Main experiment uses Llama-3.1-8B (replication with 3B confirms findings)
- Finite-sample bias: is upward-biased at finite (controlled by truncating to 1,500 tokens)
Future Directions
- Applying the Kontoyiannis cross-entropy rate to select for documents diverged from prior-generation content
- Operationalising automatic, text-only proxies for blind-writing independence in multi-agent systems
- Combining -filtering with deduplication in production pipelines
- Extending to larger model families and longer training horizons
"For iterative pipelines where model internals are inaccessible, -filtering offers a strong, zero-cost baseline."
Acknowledgments: Supported by the Australian Research Council's Discovery Projects funding scheme (DP210103700). Computational resources provided by the Phoenix HPC cluster at Adelaide University.
Related papers
- LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
LOLBENCH shows top coding agents resolve only 14% of long-horizon modular development tasks, with missing cross-module context as the dominant failure bottleneck.
- On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing
Mamba models learn linear hash functions for associative recall, requiring state memory scaling as ND = Θ(Nf log V) — linear in facts, logarithmic in vocabulary.
- Balancing Memory Pathways: Analyzing and Improving Memory Utilization in Hybrid LMs
Hybrid language models over-rely on attention and underuse recurrent memory, but an auxiliary loss forcing recurrent routing improves QA accuracy by up to 5.2% and agentic success rates.