# No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse

> The Kontoyiannis entropy rate estimator, a model-free text filter, outperforms logprob-based filtering by boosting unique trigrams 42% and cutting repetition 19% during iterative fine-tuning.

- **Source:** [arXiv](https://arxiv.org/abs/2610.01493)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/tqNGuK
- **Whiteboard:** https://picx.dev/p/tqNGuK/image

## Summary

# No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse

**Author:** Lewis Mitchell, Adelaide Data Science Centre, School of Mathematical Sciences, Adelaide University

---

## Summary (Overview)

- **Novel model-free collapse filter:** The paper introduces the Kontoyiannis entropy rate estimator $\hat{H}_K$ as a training-data filter against model collapse—the first completely model-free and reference-free filter proposed, requiring no log-probabilities, no external oracle, and no real human data.
- **Superior performance over logprob-based filtering:** In a six-generation QLoRA fine-tuning experiment on Llama-3.1-8B, $\hat{H}_K$-filtering yields **+42% unique trigrams, +30% vocabulary, and −19% repetition** (all $p < 0.001$), while logprob-based filtering ($H_L$) shows **no significant text-diversity benefit** on any metric ($p > 0.23$).
- **Validated as cross-domain entropy proxy:** $\hat{H}_K$ correlates strongly with logprob entropy ($\beta = 0.924$, $R^2 = 0.746$) across 4 domains, 2 temperatures, and 2 generator–scorer model pairs, with a domain-invariant slope.
- **Effective collapse detector:** Under fine-tuning collapse, within-topic Spearman concordance between $\hat{H}_K$ and $H_L$ is $\rho = +0.454$ ($p < 0.0001$) across 39 topics.
- **Practical and efficient:** Scoring a 1,500-word document takes under 10 ms with no GPU, making collapse-resistant curation feasible for closed-source APIs, federated training, and cross-organisation data sharing.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Large language models trained on their own outputs face **model collapse**: a self-reinforcing degradation where output entropy declines across generations, diversity narrows, and rare linguistic patterns are progressively lost (Shumailov et al., 2024; Alemohammad et al., 2024). As synthetic content becomes an increasing fraction of web-scraped training corpora, collapse mitigation is shifting from theoretical concern to engineering requirement.

### Existing Mitigation Categories

| Category | Approach | Requirement |
|----------|----------|-------------|
| **(A) Model-logprob methods** | Surplexity filtering (Gambetta et al., 2026); token resampling (Zhu et al., 2025); top-K sampling | Access to model's output distribution |
| **(B) External-oracle methods** | Verifier-based filtering (Feng et al., 2025; Yi et al., 2025) | Stronger external model or human annotator |
| **(C) Real-data accumulation** | Data accumulation (Gerstgrasser et al., 2024; Bertrand et al., 2024; Fu et al., 2025) | Continued access to human-authored data |

Each category imposes an extra burden. **Surplexity filtering**—the most established model-access-requiring baseline—requires a forward pass through the model being trained, is unavailable for closed-source APIs, computationally expensive at scale, and version-dependent. Notably, Guo et al. (2024) find that linguistic acceptability filtering can actually *worsen* collapse, underscoring that filter choice is consequential.

### The Kontoyiannis Entropy Rate Estimator

Grounded in mathematical information theory, the paper proposes the Kontoyiannis entropy rate estimator (Kontoyiannis et al., 1998):

$$
\hat{H}_K = \frac{n \log_2 n}{\sum_{i=1}^{n} \Lambda_i} \tag{1}
$$

where $\Lambda_i$ is the length of the shortest prefix of $x_i, \ldots, x_n$ that does not appear as a contiguous substring in $x_1, \ldots, x_{i-1}$. This converges almost surely to the entropy rate $H$ of any stationary ergodic process—no parametric assumptions required.

**Key properties:**
- **Match-length statistics:** Long matches signal repetition; short matches signal novelty
- **LZ compression connection:** Each $\Lambda_i$ is the elementary step of Lempel–Ziv parsing, making $\hat{H}_K$ interpretable as a normalised LZ compression rate
- **Consistency:** Grounded in the Shannon–McMillan–Breiman theorem: $-\frac{1}{n} \log_2 P(X_1, \ldots, X_n) \to H$ almost surely
- **No model, no API, no GPU required**

Prior applications include quantifying information flow in social media text (Bagrow et al., 2019; Pond et al., 2020), but no prior work has applied $\hat{H}_K$ as a collapse filter.

---

## Methodology

### Experimental Design

**Three parallel 6-generation QLoRA fine-tuning chains** on Llama-3.1-8B-Instruct:

- **80 documents generated per generation** across 4 domains (encyclopedic, creative, scientific, conversational; 10 topics × 2 documents per topic per domain)
- **40 selected for training** according to condition:
  - *Unfiltered:* random selection (seed fixed per generation)
  - *$H_L$-filter:* top 40 by mean per-token Shannon entropy (logprob-based)
  - *$\hat{H}_K$-filter:* top 40 by Kontoyiannis entropy rate (text-only)
- **Domain-stratified selection** (10 documents per domain)
- **Generation-0 documents** shared across conditions (drawn from base model)

**Training configuration:**
- 4-bit NF4 QLoRA (rank 16, α = 16)
- Learning rate $2 \times 10^{-4}$, cosine schedule, 3 epochs per generation
- Effective batch size 8
- Fixed decoding temperature of 1.0

**Tokenization:** The ProcessEntropy package's tokenizer (NLTK TweetTokenizer pass, lower-cased, non-alphanumeric characters stripped). Real-time filtering uses faster whitespace tokenization (Spearman $\rho = 0.997$ agreement).

**Finite-sample bias control:** $\hat{H}_K$ computed on first 1,500 word tokens; shorter texts discarded.

### Validation Experiments

1. **Cross-domain/cross-temperature validation:** 200 documents from GPT-4o across 4 domains, 2 temperatures ($T \in \{0.5, 1.0\}$), 25 topics per domain; 197 passed reliability filter (<5% unrecoverable logprobs)
2. **Cross-model validation:** Gemini 1.5 Pro generated additional corpora; Qwen2.5-7B-Instruct used as independent scorer
3. **Collapse detection:** Within-topic $\hat{H}_K/H_L$ concordance across generations under both rephrasing and fine-tuning regimes

---

## Empirical Validation / Results

### 1. $\hat{H}_K$ as Proxy for Logprob Entropy

**Cross-domain result:** OLS regression of $\hat{H}_K$ on $H_L$ with domain fixed effects:

$$\beta = 0.924, \quad R^2 = 0.746, \quad p < 0.0001$$

- Likelihood-ratio test for domain-specific slopes fails to reject equality ($p = 0.41$): **domain-invariant slope**
- Domains differ in baseline entropy (creative/conversational > scientific) but $\hat{H}_K$ scales with $H_L$ at the same rate in each

**Cross-model result:** Both GPT-4o/Qwen and Gemini/Qwen pairings reproduce the positive $\hat{H}_K/H_L$ relationship, confirming $\hat{H}_K$ measures a property of the text itself, not any particular model's distribution.

**Scope limitation:** On generation-0 data from Llama-3.1-8B-Instruct, correlation is near zero ($r = 0.18$, $p = 0.11$, $n = 80$). The base model produces a wide bimodal $H_L$ distribution (SD = 2.17) driven by topic-dependent familiarity rather than document-level diversity—this is the root cause of $H_L$-filtering's failure.

### 2. Collapse Detection Without Logprobs

| Regime | Within-generation concordance | Interpretation |
|--------|-------------------------------|----------------|
| Iterative rephrasing | Spearman $\rho \approx 0$ | Rephrasing homogenises surface statistics |
| Fine-tuning collapse | $\rho = +0.454$ ($p < 0.0001$) | Phrase-level repetition directly measured by $\hat{H}_K$ |

Under rephrasing, $\hat{H}_K$ and $H_L$ co-decline at indistinguishable rates ($\delta = -0.004$, $p = 0.59$), but fine-tuning collapse changes surface statistics in ways $\hat{H}_K$ directly measures.

### 3. Training Data Filter Results (Generation 6)

**Table 1: Generation-6 outcomes under three training-data conditions** (n = 80 documents per condition; p-values from Welch t-tests)

| Metric | Unfilt. | $H_L$-filt. | $\hat{H}_K$-filt. | $p(\hat{H}_K \text{ v. U})$ | $p(H_L \text{ v. U})$ | $p(\hat{H}_K \text{ v. } H_L)$ |
|--------|---------|-------------|-------------------|---------------------------|----------------------|-------------------------------|
| $\hat{H}_K$ (bits/word) | 0.79 | 1.16 | 1.43 | <0.001 | 0.008 | <0.001 |
| Distinct-3 (%) | — | +7% | +42% | <0.001 | 0.23 | <0.001 |
| Vocabulary (%) | — | +4% | +30% | <0.001 | 0.31 | <0.001 |
| Rep-4 (%) | — | −3% | −19% | <0.001 | 0.27 | <0.001 |

**Key findings:**

- **$\hat{H}_K$-filtering** retains 1.43 bits/word vs. 0.79 unfiltered ($p < 0.001$)
- **$H_L$-filtering provides no text-diversity benefit** ($p > 0.23$ on all diversity metrics). Mechanistic explanation: $H_L$ distribution compresses dramatically after one fine-tuning step (gen-0 SD = 2.17 → gen-1 SD = 0.13), collapsing the ranking signal
- **Filter selection sets differ substantially:** within-domain Jaccard similarity ≈ 0.35 (range 0.18–0.54), indicating complementary axes

### 4. Comparison with Simpler Text-Only Baselines

Static-pool comparison on generation-0 documents (exact comparison, all conditions share same 80 documents):

| Filter | Mean $\hat{H}_K$ of selected docs | Jaccard vs. $\hat{H}_K$-filter |
|--------|-----------------------------------|-------------------------------|
| $\hat{H}_K$-filtering | 3.98 | — |
| TTR@500-filtering | 3.37 | 0.43 |
| Distinct-2-filtering | 3.32 | — |
| gzip-compression ratio | — | 0.36–0.38 |
| Repetition-rate filter | — | 0.36–0.38 |

As genuine six-generation training filters, $\hat{H}_K$-filtering significantly outperforms both gzip-filtering (Distinct-3 +17%, vocabulary +17%, Rep-4 −7%; all $p < 0.03$) and repetition-rate filtering (no benefit; all $p > 0.24$) on every metric (all $p < 0.005$, most $p < 10^{-4}$).

### 5. Downstream Benchmarks

- **HellaSwag:** Accuracy drops from 0.729 (gen-0) to ≈0.652–0.657 (gen-6) with no significant difference between conditions—$\hat{H}_K$-filtering preserves text diversity but not general reasoning ability
- **MMLU (5-shot):** 67.6–67.6% across conditions at gen-6 (all pairwise $\chi^2$ tests $p > 0.8$)
- **GSM8K (8-shot):** 70.7–71.2% across conditions at gen-6 (all $p > 0.8$)
- **MAUVE scores:** Near-ceiling and uniform (0.987–0.993), indicating embedding-space similarity preserved

### 6. Quality Assessment

LLM-judge pass (GPT-4o-mini, blind to condition) on all 240 generation-6 documents:

| Metric | $\hat{H}_K$-filt. | Unfiltered | $H_L$-filt. |
|--------|-------------------|------------|-------------|
| Coherence (1–5) | 2.35 | 2.20 ($p = 0.058$) | 2.29 |
| Instruction-following (1–5) | 3.23 | 3.00 ($p = 0.020$) | 3.15 |

$\hat{H}_K$-filtering's diversity advantage comes with a small quality benefit rather than a coherence cost.

### 7. Replication with Smaller Model

With Llama-3.2-3B-Instruct (same corpus and protocol):
- $\hat{H}_K$-filtering retains 2.91 bits/word vs. 2.12 unfiltered ($p < 0.001$) and 2.41 for $H_L$-filtering ($p = 0.004$)
- $H_L$-filtering again non-significant on all text-diversity metrics
- $\hat{H}_K$-vs-$H_L$ comparison reaches significance on the $\hat{H}_K$ metric ($p = 0.004$ vs. $p = 0.088$ at 8B)

---

## Theoretical and Practical Implications

### Why $\hat{H}_K$-filtering outperforms $H_L$-filtering

1. **Direct targeting of collapse signature:** $\hat{H}_K$ directly penalises repeated phrase structure—the primary surface manifestation of fine-tuning collapse. Documents with recurring n-grams exhibit long match lengths and hence low $\hat{H}_K$.

2. **$H_L$ signal destruction:** $H_L$ measures token-level uncertainty in the scorer model—a signal destroyed after a single fine-tuning step. At generation-0, $H_L$ primarily reflects topic familiarity rather than document diversity.

3. **Surplexity collapse:** A literal reimplementation shows median absolute deviation falls from 4.40 (gen-0 pool) to 0.40–0.49 after one fine-tuning step—a 9–11-fold narrowing.

### Simpson's Paradox and the Jaccard Gap

The two filters select different documents (Jaccard ≈ 0.35) despite corpus-level correlation ($R^2 = 0.746$) due to a form of Simpson's paradox:
- **Between-domain correlation is strong** (creative/conversational have both high $H_L$ and high $\hat{H}_K$)
- **Within-domain residual correlation is near zero**—where filtering operates due to domain stratification

The estimators measure genuinely different properties of intra-domain text variation.

### Practical Implications

- **Zero-cost filtering:** Scoring a 1,500-word document takes under 10 ms; a pipeline producing $10^5$ candidate documents per round incurs negligible, trivially parallelisable overhead
- **Applicable where logprob access is unavailable:** closed-source API distillation, cross-organisation data sharing, federated training
- **Complementary to deduplication:** Min-hash near-deduplication removes between-document duplicates; $\hat{H}_K$-filtering ensures within-document structure hasn't collapsed. Combined pipeline (deduplicate first, then $\hat{H}_K$-filter) addresses both failure modes
- **Multi-agent diversity:** The Kontoyiannis cross-entropy rate $\hat{h}_\times(T \mid S)$ could operationalise the blind-writing independence phase Chen et al. (2026) identify as the most effective structural intervention against diversity collapse

---

## Conclusion

### Main Takeaways

The Kontoyiannis entropy rate estimator $\hat{H}_K$—a model-free, match-length measure of sequential complexity—is an effective training-data filter against fine-tuning-driven model collapse:

- **+42% unique trigrams, +30% vocabulary, −19% intra-document repetition** vs. unfiltered training at generation 6
- **Logprob-based filtering shows no significant diversity benefit** ($p > 0.23$ on all metrics)
- **Zero-cost, model-free, reference-free:** no model access, no API calls, no GPU required

### Key Limitations

1. **Scope:** $\hat{H}_K$-filtering preserves text diversity but does not protect general reasoning ability (HellaSwag, MMLU, GSM8K all degrade uniformly)
2. **Specificity:** Advantage is specific to fine-tuning collapse, not iterative rephrasing (where within-generation rank concordance is ≈ 0)
3. **Lexical focus:** Diversity gains are lexical (MAUVE scores near-ceiling and uniform across conditions)
4. **Single model family:** Main experiment uses Llama-3.1-8B (replication with 3B confirms findings)
5. **Finite-sample bias:** $\hat{H}_K$ is upward-biased at finite $n$ (controlled by truncating to 1,500 tokens)

### Future Directions

- Applying the Kontoyiannis cross-entropy rate $\hat{h}_\times(T \mid S)$ to select for documents diverged from prior-generation content
- Operationalising automatic, text-only proxies for blind-writing independence in multi-agent systems
- Combining $\hat{H}_K$-filtering with deduplication in production pipelines
- Extending to larger model families and longer training horizons

> *"For iterative pipelines where model internals are inaccessible, $\hat{H}_K$-filtering offers a strong, zero-cost baseline."*

---

**Acknowledgments:** Supported by the Australian Research Council's Discovery Projects funding scheme (DP210103700). Computational resources provided by the Phoenix HPC cluster at Adelaide University.

---

_Markdown view of https://picx.dev/p/tqNGuK, served by PicX — AI-generated visual whiteboard summaries of research papers._
