Skip to main content
Back to timeline
arXivSource publication:

No Model Required: Text Entropy Rate Filtering Yields +42% Unique Trigrams, +30% Vocabulary, and -19% Repetition in Six-Generation QLoRA Iterative Fine-Tuning

Synopsis

The work introduces a non-parametric Kontoyiannis entropy rate estimator h_k, computed entirely from raw text with no model of any kind, as a training-data filter; in a six-generation QLoRA collapse experiment on Llama-3.1-8B, h_k-filtering yields +42% unique trigrams, +30% vocabulary, and -19% repetition (all p<0.001), whereas logprob-based filtering provides no significant text-diversity benefit on any metric (p>0.23); h_k is further validated as a cross-domain entropy proxy (beta=0.924, R^2=0.746) and collapse detector (rho=+0.454, p<0.0001) across 4 domains, 2 temperatures, 2 generator-scorer model pairs, and 1,520 generated documents.

Source-provided article image: No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse
Figure 1 ·

Figure 1 : Filtering on text-only entropy preserves diversity; logprob-based filtering does not. (a) Kontoyiannis entropy rate H ^ K \hat{H}_{K} across six fine-tuning generations of Llama-3.1-8B-Instruct under three training-data conditions: unfiltered (rapid collapse), H L H_{L} -filtered (intermediate), and H ^ K \hat{H}_{K} -filtered (slowest decline). Shaded bands: ± 1 \pm 1 SEM ( n = 80 n=80 documents per generation). (b) Percentage change relative to the unfiltered baseline at generation 6 for three text-diversity metrics (Distinct-3, Vocabulary size, Rep-4). Error bars: bootstrap 95% CIs. H ^ K \hat{H}_{K} -based filtering yields highly significant improvements on all three metrics ( ∗ ∗ ∗ p < 0.001 {}^{***}p<0.001 ), whereas H L H_{L} -based filtering is non-significant on all three.

arXiv

Interpretation

A non-parametric Kontoyiannis entropy rate estimator h_k, computed entirely from raw text via match-length statistics with no model of any kind, is proposed and validated as a training-data filter that mitigates model collapse in iterative fine-tuning. Existing mitigations require model log-probabilities, an external oracle, or continued access to real human data; this work shows that filtering based solely on raw-text entropy rate is feasible and, in a fully-synthetic single-lineage setting, outperforms the most established model-access-requiring logprob-based baseline. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, h_k-filtering yields +42% unique trigrams, +30% vocabulary, and -19% repetition (all p<0.001), while logprob-based filtering provides no significant benefit on any metric (p>0.23).

h_k serves as a cross-domain entropy proxy and collapse detector, maintaining stable associations across diverse generation conditions. Extends an information-theoretic entropy rate estimator from a theoretical tool to a cross-domain text-diversity proxy and collapse signal, covering 4 domains, 2 temperatures, 2 generator-scorer model pairs, and 1,520 generated documents. Cross-domain entropy proxy regression yields beta=0.924 and R^2=0.746; collapse detection correlation yields rho=+0.454, p<0.0001.

The results demonstrate that information-theoretic approaches to collapse mitigation are efficient and suggest new approaches for maintaining multi-agent diversity. Shifts collapse mitigation away from reliance on model-internal signals toward pure text statistics, reducing dependence on external oracles, model log-probabilities, or continued real data. Based on the six-generation QLoRA experiment and validation across 4 domains, 2 temperatures, 2 model pairs, and 1,520 documents.

Perspective

The result applies to fully-synthetic, single-lineage iterative fine-tuning settings, where no real human data flows back and only the model's own generated data is used repeatedly; in this setting h_k-filtering offers a lightweight, model-access-free data selection mechanism. For teams that need to maintain text diversity in a training pipeline but cannot obtain model log-probabilities or an external oracle, this method provides an actionable alternative; its cross-domain entropy proxy and collapse detection properties also suit diversity monitoring and early collapse warning on generated documents.

Based only on the abstract, it is not possible to confirm the specific match-length parameter of the h_k estimator, the filtering threshold selection, computational cost, or the per-generation sample size and data composition in the six-generation experiment; the names of the 4 domains, the 2 temperature values, and the identities of the 2 generator-scorer model pairs are also not given in the abstract. In addition, the abstract does not state how the method performs under mixed real-data or non-single-lineage settings, which are open questions a reader should verify before adoption.

Sources