Implicit Bias in Hyperbolic Multiclass Learning Is Governed by Busemann Risk: A Drift Coefficient Decides Whether Embeddings Escape to the Ideal Boundary or Return Inward
Synopsis
Studying the implicit bias of Riemannian gradient flow for hyperbolic multiclass classification with fixed prototypes, this work decomposes large-radius distances into a radial term and a Busemann direction term, proving a radial dichotomy in which the sign of a drift coefficient decides whether the radius escapes to the ideal boundary or returns inward, and showing that boundary directions converge to critical points of the Busemann risk, with an asymptotic classifier and decision boundaries on the ideal boundary.
(a) Hyperboloid model and Poincaré Disk
arXivInterpretation
It introduces a radial Busemann expansion: at large radius, the hyperbolic distance from an embedding to any fixed prototype splits into a common radial term and a direction-dependent term given by the Busemann function. Prior hyperbolic representation learning work focused mainly on model design and empirical phenomena, lacking an asymptotic account of long-time training dynamics; this expansion separates the nonconvex loss into radial and angular parts, enabling the later convergence analysis. The decomposition is stated as a lemma with a full proof, the remainder is explicitly bounded as exponentially small, and the threshold depends only on the prototype, not on the direction.
It proves a radial dichotomy: the sign of the drift coefficient determines whether the radius is pushed toward the ideal boundary or pulled back toward the interior; persistent positive drift makes the radius diverge at a logarithmic rate, while persistent negative drift returns the trajectory to the large-radius threshold in finite time. In Euclidean separable linear classification the logarithmic growth coefficient is set by the data margin, whereas here the drift coefficient is universal in the studied setting, independent of data geometry or margin, giving a mechanism for boundary saturation in hyperbolic models. The result is a theorem under the totally regular PERM loss assumption, with upper and lower bounds on the escape rate and a bound on the first hitting time in the negative-drift case, supported by controlled synthetic experiments fitting the predicted slope.
It proves angular convergence: under a logarithmic radial tail condition, the boundary direction of an escaping trajectory converges to a critical point of the Busemann risk on the ideal boundary. This offers a mechanism for near-boundary clustering: once embeddings have escaped sufficiently, discriminative information is encoded mainly in boundary directions governed by the Busemann risk rather than in radial positions. The proof invokes the Chill and Jendoubi convergence framework for asymptotically autonomous gradient flows, verifying regularity, precompactness, the Łojasiewicz gradient inequality, and perturbation decay.
It characterizes the asymptotic classifier: when the minimizer of the Busemann risk is unique, prediction at large radius is decided by the smallest Busemann value, and the decision boundary between classes is the intersection of the ideal boundary with an affine hyperplane, i.e., a spherical section. This reduces the late-stage behavior of hyperbolic classification to geometric objects on the ideal boundary and defines a Busemann margin, making correct classification equivalent to that margin being positive. The result is a theorem under the totally regular PERM loss assumption, proved using uniform error bounds from the radial Busemann expansion together with permutation invariance and strict coordinate-wise decrease of the template.
Perspective
The results apply to hyperbolic multiclass classification with fixed class prototypes, distance-based scores, and totally regular PERM losses, and to the training stage after embeddings have entered the large-radius asymptotic regime. For practitioners, late-stage classification behavior is reduced to the Busemann risk on the ideal boundary, so radius clipping can be read as preventing excessive radial growth while preserving the boundary direction structure that determines the classifier; a curvature-scaling corollary shows the conclusions rescale with the Minkowski constant. For theorists, the framework can be transferred to other relative-margin multiclass losses and applies trajectory-wise, including individual samples in nonseparable settings.
Angular convergence relies on a logarithmic radial tail condition, which excludes degenerate regimes where the leading radial drift vanishes and higher-order terms set the escape rate; such borderline cases remain open. The analysis fixes prototypes, so additional interactions from jointly learning embeddings, prototypes, and encoder parameters in the fully coupled setting are not covered. Experiments are controlled synthetic verifications rather than large-scale benchmarks, so variation from encoder architecture, stochastic optimization, finite-sample noise, and model misspecification on real data remains to be examined. In addition, some formulas and numerical values are absent from the parsed text, so checking specific constants and integrator settings requires the original appendix.
