Re-bounding the Human-Data Ratio Needed to Prevent Model Collapse via Fisher-Rao Geometry
Synopsis
This work models iterative generative-model training as a closed-loop stochastic process on the probability simplex and analyzes its dynamics under the Fisher-Rao metric instead of the Euclidean metric, deriving contraction and invariance bounds that remain meaningful as dimension grows and giving a human-to-synthetic data ratio threshold that guarantees convergence to a Fisher-Rao ball, concluding that the effective required human-data ratio is higher than previously implied.
Interpretation
It provides quantitative contraction and invariance bounds for the iterative-training dynamics under the Fisher-Rao metric, together with an explicit threshold condition on the human-to-synthetic ratio μ. Prior work [15] gave convergence to a Euclidean ball under the Euclidean metric; this paper works directly on the information-geometric structure of the probability simplex, yielding the Fisher-Rao ball bound in Eq. (8), the ball radius ϵ_FR in Eq. (9), and the μ threshold in Eq. (7). The result is a rigorous theorem-form derivation relying on Assumption 1 (bounded perturbation ‖ε(t)‖∞≤η and strictly positive human distribution δ>0) and Assumption 2 (existence of an equilibrium θ_e), proved step by step through Lemmas 1–8; it is a mathematical proof rather than an experimental validation.
Euclidean bounds become vacuous in high dimensions, whereas the Fisher-Rao bound does not lose meaning as dimension grows. The paper notes that as n→∞ the Euclidean distance between almost any two probability distributions approaches zero, and gives an example of “disjoint” probability vectors p and q whose Euclidean distance scales as O(1/√n) while their Fisher-Rao distance stays constant at π/2. This contrast is supported by the explicit construction and distance computation given in the text; it is a conceptual counterexample argument motivating the change of metric.
Under stated scaling laws the two metrics give different decay rates of the error bound with dimension n, and the Fisher-Rao analysis requires the human-data ratio to grow faster. Proposition 1, under δ_n∼n^(−β0), μ_n∼c n^p, p>2β0, gives ϵ=Θ(n^(−p)) in the Euclidean case and ϵ_FR=Θ(n^(1/2+β0−p)) in the Fisher-Rao case. The scaling conclusion comes from the proof of Proposition 1, which includes bounds on how θ_e varies with n (e.g., ‖θ_e−θ_0‖∞=O(n^(−p)) and min_i θ_e,i=Θ(n^(−β0))).
To obtain, for example, O(1/n)-order decay, a natural choice of human-data ratio is μ_n=n^(5/2+γ) with γ>0, whereas the Euclidean estimate predicts 1/n-order error already at μ_n∼n. The paper argues on this basis that the Euclidean estimate does not account for the geometric cost of spreading probability mass across n coordinates, and that μ_n∼n does not prevent the Fisher-Rao error from growing, so stronger growth of μ_n is required. This statement is an illustrative choice made without extra structural information on θ_e, based on the scaling relations of Proposition 1; it is a theoretical corollary rather than an experimental measurement.
Perspective
The result applies to iterative-training settings modeled as the closed-loop stochastic process in Eq. (5), requiring a bounded perturbation, a strictly positive human distribution, and the existence of an equilibrium θ_e, and assuming a constant human-to-synthetic ratio μ=β_k/α_k; under these conditions it gives the size of the Fisher-Rao ball, the convergence rate, and the μ threshold guaranteeing this behavior, which can guide lower-bound estimates for mixing real and synthetic data.
A careful reader would still watch: how closely the continuous-time limit and constant-ratio μ assumption match real discrete, batched training pipelines; how the quantities δ, η, T, and θ_e in the threshold Eq. (7) and radius Eq. (9) can be estimated in actual models; and how feasible choices like μ_n=n^(5/2+γ) are for concrete tasks when θ_e has no structure beyond the simplex constraints. This is a theoretical analysis without experimental validation, so the practical values and effects of these quantities remain open questions.
