No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse

Author: Lewis Mitchell, Adelaide Data Science Centre, School of Mathematical Sciences, Adelaide University


Summary (Overview)

  • Novel model-free collapse filter: The paper introduces the Kontoyiannis entropy rate estimator H^K\hat{H}_K as a training-data filter against model collapse—the first completely model-free and reference-free filter proposed, requiring no log-probabilities, no external oracle, and no real human data.
  • Superior performance over logprob-based filtering: In a six-generation QLoRA fine-tuning experiment on Llama-3.1-8B, H^K\hat{H}_K-filtering yields +42% unique trigrams, +30% vocabulary, and −19% repetition (all p<0.001p < 0.001), while logprob-based filtering (HLH_L) shows no significant text-diversity benefit on any metric (p>0.23p > 0.23).
  • Validated as cross-domain entropy proxy: H^K\hat{H}_K correlates strongly with logprob entropy (β=0.924\beta = 0.924, R2=0.746R^2 = 0.746) across 4 domains, 2 temperatures, and 2 generator–scorer model pairs, with a domain-invariant slope.
  • Effective collapse detector: Under fine-tuning collapse, within-topic Spearman concordance between H^K\hat{H}_K and HLH_L is ρ=+0.454\rho = +0.454 (p<0.0001p < 0.0001) across 39 topics.
  • Practical and efficient: Scoring a 1,500-word document takes under 10 ms with no GPU, making collapse-resistant curation feasible for closed-source APIs, federated training, and cross-organisation data sharing.

Introduction and Theoretical Foundation

Background and Motivation

Large language models trained on their own outputs face model collapse: a self-reinforcing degradation where output entropy declines across generations, diversity narrows, and rare linguistic patterns are progressively lost (Shumailov et al., 2024; Alemohammad et al., 2024). As synthetic content becomes an increasing fraction of web-scraped training corpora, collapse mitigation is shifting from theoretical concern to engineering requirement.

Existing Mitigation Categories

CategoryApproachRequirement
(A) Model-logprob methodsSurplexity filtering (Gambetta et al., 2026); token resampling (Zhu et al., 2025); top-K samplingAccess to model's output distribution
(B) External-oracle methodsVerifier-based filtering (Feng et al., 2025; Yi et al., 2025)Stronger external model or human annotator
(C) Real-data accumulationData accumulation (Gerstgrasser et al., 2024; Bertrand et al., 2024; Fu et al., 2025)Continued access to human-authored data

Each category imposes an extra burden. Surplexity filtering—the most established model-access-requiring baseline—requires a forward pass through the model being trained, is unavailable for closed-source APIs, computationally expensive at scale, and version-dependent. Notably, Guo et al. (2024) find that linguistic acceptability filtering can actually worsen collapse, underscoring that filter choice is consequential.

The Kontoyiannis Entropy Rate Estimator

Grounded in mathematical information theory, the paper proposes the Kontoyiannis entropy rate estimator (Kontoyiannis et al., 1998):

H^K=nlog⁡2n∑i=1nΛi(1)\hat{H}_K = \frac{n \log_2 n}{\sum_{i=1}^{n} \Lambda_i} \tag{1}

where Λi\Lambda_i is the length of the shortest prefix of xi,…,xnx_i, \ldots, x_n that does not appear as a contiguous substring in x1,…,xi−1x_1, \ldots, x_{i-1}. This converges almost surely to the entropy rate HH of any stationary ergodic process—no parametric assumptions required.

Key properties:

  • Match-length statistics: Long matches signal repetition; short matches signal novelty
  • LZ compression connection: Each Λi\Lambda_i is the elementary step of Lempel–Ziv parsing, making H^K\hat{H}_K interpretable as a normalised LZ compression rate
  • Consistency: Grounded in the Shannon–McMillan–Breiman theorem: −1nlog⁡2P(X1,…,Xn)→H-\frac{1}{n} \log_2 P(X_1, \ldots, X_n) \to H almost surely
  • No model, no API, no GPU required

Prior applications include quantifying information flow in social media text (Bagrow et al., 2019; Pond et al., 2020), but no prior work has applied H^K\hat{H}_K as a collapse filter.


Methodology

Experimental Design

Three parallel 6-generation QLoRA fine-tuning chains on Llama-3.1-8B-Instruct:

  • 80 documents generated per generation across 4 domains (encyclopedic, creative, scientific, conversational; 10 topics × 2 documents per topic per domain)
  • 40 selected for training according to condition:
    • Unfiltered: random selection (seed fixed per generation)
    • HLH_L-filter: top 40 by mean per-token Shannon entropy (logprob-based)
    • H^K\hat{H}_K-filter: top 40 by Kontoyiannis entropy rate (text-only)
  • Domain-stratified selection (10 documents per domain)
  • Generation-0 documents shared across conditions (drawn from base model)

Training configuration:

  • 4-bit NF4 QLoRA (rank 16, α = 16)
  • Learning rate 2×10−42 \times 10^{-4}, cosine schedule, 3 epochs per generation
  • Effective batch size 8
  • Fixed decoding temperature of 1.0

Tokenization: The ProcessEntropy package's tokenizer (NLTK TweetTokenizer pass, lower-cased, non-alphanumeric characters stripped). Real-time filtering uses faster whitespace tokenization (Spearman ρ=0.997\rho = 0.997 agreement).

Finite-sample bias control: H^K\hat{H}_K computed on first 1,500 word tokens; shorter texts discarded.

Validation Experiments

  1. Cross-domain/cross-temperature validation: 200 documents from GPT-4o across 4 domains, 2 temperatures (T∈{0.5,1.0}T \in \{0.5, 1.0\}), 25 topics per domain; 197 passed reliability filter (<5% unrecoverable logprobs)
  2. Cross-model validation: Gemini 1.5 Pro generated additional corpora; Qwen2.5-7B-Instruct used as independent scorer
  3. Collapse detection: Within-topic H^K/HL\hat{H}_K/H_L concordance across generations under both rephrasing and fine-tuning regimes

Empirical Validation / Results

1. H^K\hat{H}_K as Proxy for Logprob Entropy

Cross-domain result: OLS regression of H^K\hat{H}_K on HLH_L with domain fixed effects:

β=0.924,R2=0.746,p<0.0001\beta = 0.924, \quad R^2 = 0.746, \quad p < 0.0001
  • Likelihood-ratio test for domain-specific slopes fails to reject equality (p=0.41p = 0.41): domain-invariant slope
  • Domains differ in baseline entropy (creative/conversational > scientific) but H^K\hat{H}_K scales with HLH_L at the same rate in each

Cross-model result: Both GPT-4o/Qwen and Gemini/Qwen pairings reproduce the positive H^K/HL\hat{H}_K/H_L relationship, confirming H^K\hat{H}_K measures a property of the text itself, not any particular model's distribution.

Scope limitation: On generation-0 data from Llama-3.1-8B-Instruct, correlation is near zero (r=0.18r = 0.18, p=0.11p = 0.11, n=80n = 80). The base model produces a wide bimodal HLH_L distribution (SD = 2.17) driven by topic-dependent familiarity rather than document-level diversity—this is the root cause of HLH_L-filtering's failure.

2. Collapse Detection Without Logprobs

RegimeWithin-generation concordanceInterpretation
Iterative rephrasingSpearman ρ≈0\rho \approx 0Rephrasing homogenises surface statistics
Fine-tuning collapseρ=+0.454\rho = +0.454 (p<0.0001p < 0.0001)Phrase-level repetition directly measured by H^K\hat{H}_K

Under rephrasing, H^K\hat{H}_K and HLH_L co-decline at indistinguishable rates (δ=−0.004\delta = -0.004, p=0.59p = 0.59), but fine-tuning collapse changes surface statistics in ways H^K\hat{H}_K directly measures.

3. Training Data Filter Results (Generation 6)

Table 1: Generation-6 outcomes under three training-data conditions (n = 80 documents per condition; p-values from Welch t-tests)

MetricUnfilt.HLH_L-filt.H^K\hat{H}_K-filt.p(H^K v. U)p(\hat{H}_K \text{ v. U})p(HL v. U)p(H_L \text{ v. U})p(H^K v. HL)p(\hat{H}_K \text{ v. } H_L)
H^K\hat{H}_K (bits/word)0.791.161.43<0.0010.008<0.001
Distinct-3 (%)—+7%+42%<0.0010.23<0.001
Vocabulary (%)—+4%+30%<0.0010.31<0.001
Rep-4 (%)—−3%−19%<0.0010.27<0.001

Key findings:

  • H^K\hat{H}_K-filtering retains 1.43 bits/word vs. 0.79 unfiltered (p<0.001p < 0.001)
  • HLH_L-filtering provides no text-diversity benefit (p>0.23p > 0.23 on all diversity metrics). Mechanistic explanation: HLH_L distribution compresses dramatically after one fine-tuning step (gen-0 SD = 2.17 → gen-1 SD = 0.13), collapsing the ranking signal
  • Filter selection sets differ substantially: within-domain Jaccard similarity ≈ 0.35 (range 0.18–0.54), indicating complementary axes

4. Comparison with Simpler Text-Only Baselines

Static-pool comparison on generation-0 documents (exact comparison, all conditions share same 80 documents):

FilterMean H^K\hat{H}_K of selected docsJaccard vs. H^K\hat{H}_K-filter
H^K\hat{H}_K-filtering3.98—
TTR@500-filtering3.370.43
Distinct-2-filtering3.32—
gzip-compression ratio—0.36–0.38
Repetition-rate filter—0.36–0.38

As genuine six-generation training filters, H^K\hat{H}_K-filtering significantly outperforms both gzip-filtering (Distinct-3 +17%, vocabulary +17%, Rep-4 −7%; all p<0.03p < 0.03) and repetition-rate filtering (no benefit; all p>0.24p > 0.24) on every metric (all p<0.005p < 0.005, most p<10−4p < 10^{-4}).

5. Downstream Benchmarks

  • HellaSwag: Accuracy drops from 0.729 (gen-0) to ≈0.652–0.657 (gen-6) with no significant difference between conditions—H^K\hat{H}_K-filtering preserves text diversity but not general reasoning ability
  • MMLU (5-shot): 67.6–67.6% across conditions at gen-6 (all pairwise χ2\chi^2 tests p>0.8p > 0.8)
  • GSM8K (8-shot): 70.7–71.2% across conditions at gen-6 (all p>0.8p > 0.8)
  • MAUVE scores: Near-ceiling and uniform (0.987–0.993), indicating embedding-space similarity preserved

6. Quality Assessment

LLM-judge pass (GPT-4o-mini, blind to condition) on all 240 generation-6 documents:

MetricH^K\hat{H}_K-filt.UnfilteredHLH_L-filt.
Coherence (1–5)2.352.20 (p=0.058p = 0.058)2.29
Instruction-following (1–5)3.233.00 (p=0.020p = 0.020)3.15

H^K\hat{H}_K-filtering's diversity advantage comes with a small quality benefit rather than a coherence cost.

7. Replication with Smaller Model

With Llama-3.2-3B-Instruct (same corpus and protocol):

  • H^K\hat{H}_K-filtering retains 2.91 bits/word vs. 2.12 unfiltered (p<0.001p < 0.001) and 2.41 for HLH_L-filtering (p=0.004p = 0.004)
  • HLH_L-filtering again non-significant on all text-diversity metrics
  • H^K\hat{H}_K-vs-HLH_L comparison reaches significance on the H^K\hat{H}_K metric (p=0.004p = 0.004 vs. p=0.088p = 0.088 at 8B)

Theoretical and Practical Implications

Why H^K\hat{H}_K-filtering outperforms HLH_L-filtering

  1. Direct targeting of collapse signature: H^K\hat{H}_K directly penalises repeated phrase structure—the primary surface manifestation of fine-tuning collapse. Documents with recurring n-grams exhibit long match lengths and hence low H^K\hat{H}_K.

  2. HLH_L signal destruction: HLH_L measures token-level uncertainty in the scorer model—a signal destroyed after a single fine-tuning step. At generation-0, HLH_L primarily reflects topic familiarity rather than document diversity.

  3. Surplexity collapse: A literal reimplementation shows median absolute deviation falls from 4.40 (gen-0 pool) to 0.40–0.49 after one fine-tuning step—a 9–11-fold narrowing.

Simpson's Paradox and the Jaccard Gap

The two filters select different documents (Jaccard ≈ 0.35) despite corpus-level correlation (R2=0.746R^2 = 0.746) due to a form of Simpson's paradox:

  • Between-domain correlation is strong (creative/conversational have both high HLH_L and high H^K\hat{H}_K)
  • Within-domain residual correlation is near zero—where filtering operates due to domain stratification

The estimators measure genuinely different properties of intra-domain text variation.

Practical Implications

  • Zero-cost filtering: Scoring a 1,500-word document takes under 10 ms; a pipeline producing 10510^5 candidate documents per round incurs negligible, trivially parallelisable overhead
  • Applicable where logprob access is unavailable: closed-source API distillation, cross-organisation data sharing, federated training
  • Complementary to deduplication: Min-hash near-deduplication removes between-document duplicates; H^K\hat{H}_K-filtering ensures within-document structure hasn't collapsed. Combined pipeline (deduplicate first, then H^K\hat{H}_K-filter) addresses both failure modes
  • Multi-agent diversity: The Kontoyiannis cross-entropy rate h^×(T∣S)\hat{h}_\times(T \mid S) could operationalise the blind-writing independence phase Chen et al. (2026) identify as the most effective structural intervention against diversity collapse

Conclusion

Main Takeaways

The Kontoyiannis entropy rate estimator H^K\hat{H}_K—a model-free, match-length measure of sequential complexity—is an effective training-data filter against fine-tuning-driven model collapse:

  • +42% unique trigrams, +30% vocabulary, −19% intra-document repetition vs. unfiltered training at generation 6
  • Logprob-based filtering shows no significant diversity benefit (p>0.23p > 0.23 on all metrics)
  • Zero-cost, model-free, reference-free: no model access, no API calls, no GPU required

Key Limitations

  1. Scope: H^K\hat{H}_K-filtering preserves text diversity but does not protect general reasoning ability (HellaSwag, MMLU, GSM8K all degrade uniformly)
  2. Specificity: Advantage is specific to fine-tuning collapse, not iterative rephrasing (where within-generation rank concordance is ≈ 0)
  3. Lexical focus: Diversity gains are lexical (MAUVE scores near-ceiling and uniform across conditions)
  4. Single model family: Main experiment uses Llama-3.1-8B (replication with 3B confirms findings)
  5. Finite-sample bias: H^K\hat{H}_K is upward-biased at finite nn (controlled by truncating to 1,500 tokens)

Future Directions

  • Applying the Kontoyiannis cross-entropy rate h^×(T∣S)\hat{h}_\times(T \mid S) to select for documents diverged from prior-generation content
  • Operationalising automatic, text-only proxies for blind-writing independence in multi-agent systems
  • Combining H^K\hat{H}_K-filtering with deduplication in production pipelines
  • Extending to larger model families and longer training horizons

"For iterative pipelines where model internals are inaccessible, H^K\hat{H}_K-filtering offers a strong, zero-cost baseline."


Acknowledgments: Supported by the Australian Research Council's Discovery Projects funding scheme (DP210103700). Computational resources provided by the Phoenix HPC cluster at Adelaide University.

Related papers