The Numerical Stability of Hyperbolic Representation Learning
Gal MishneZhengchao WanYusu WangSheng Yang
Analyzes the representation and optimization trade-offs between Poincaré and Lorentz models under standard floating-point precision, presenting a Euclidean parametrization that prevents numerical instabilities in hyperbolic machine learning.
Hyperbolic machine learning has become an essential technique for analyzing hierarchical and tree-like data across domains such as natural language processing, recommendation systems, and genomics. Hyperbolic space naturally accommodates branching structures with low distortion because its volume expands exponentially. However, this same geometric property creates severe numerical instability during computational modeling. In standard 64-bit floating-point arithmetic, existing implementations routinely encounter vanishing gradients, rounding limits, and invalid calculations (NaN errors), forcing systems to constrain learning parameters or halt entirely.
The article systematically evaluates the numerical representation limits and optimization dynamics of the two most common hyperbolic frameworks—the Poincaré ball and the Lorentz model. It demonstrates mathematically and empirically why standard models fail, introduces a robust Euclidean parametrization to resolve these trade-offs, and extends this framework to hyperbolic support vector machines for classification tasks.
To conduct this evaluation, the authors performed theoretical mathematical analyses of coordinate capacities and gradient behaviors in 64-bit precision. They supplemented this theory with empirical validation on eight simulated tree-embedding datasets and six real-world classification benchmarks, including image datasets (CIFAR-10, Fashion-MNIST) and multiple single-cell RNA biological datasets. Performance was evaluated across metric distortion, gradient norms, embedding spread, classification accuracy, and macro F1 scores.
The findings establish that under 64-bit arithmetic, the Poincaré model has a wider representation capacity radius of approximately 38 from the origin before points collapse to the boundary, compared to a radius of approximately 19 for the Lorentz model. Conversely, the Lorentz model is substantially superior in optimization; the Poincaré model suffers from severe gradient vanishing near the boundary because its gradient update terms scale at second-order decay (10^-2k) versus first-order decay (10^-k) in Lorentz. The proposed Euclidean feature parametrization successfully combines the benefits of both by eliminating numerical capacity boundaries entirely while maintaining first-order optimization dynamics comparable to Lorentz. In empirical tree embeddings, the Euclidean approach achieved lower average distortion (reaching 1.0019–1.0292) and substantially larger embedding diameters (around 8.8–19.8) compared to the restricted Poincaré embeddings (around 4.2–6.4). Finally, applying this parametrization to hyperbolic hyperplanes removed non-convex constraints, creating a reformulated classifier (LSVMPP) that outperformed standard Euclidean, Poincaré, and Lorentz baselines across almost all benchmarks (e.g., reaching 89.49% on Fashion-MNIST and 62.64% on the Paul dataset).
These results demonstrate that practitioner difficulties with hyperbolic learning stem from structural flaws in coordinate optimization rather than data quality. By utilizing Euclidean parametrizations, engineering teams can eliminate catastrophic training collapses, avoid artificial thresholding hacks, and improve model accuracy on complex hierarchical data without incurring custom ultra-high-precision hardware overhead.
Engineering and research teams should adopt Euclidean parametrizations as the default training proxy for hyperbolic representation learning and hyperbolic support vector machines. Additionally, practitioners employing multi-class scaling should use hyperbolic signed distance transformations (arcsinh adjustments) to calibrate prediction confidence effectively. Before broad deployment, development teams should conduct pilots on larger-scale networks, as accelerated momentum methods in hyperbolic spaces still require further investigation.
The article's conclusions are supported with high confidence by both formal proofs and consistent multi-domain experiments. However, practitioners should note that computing certain derived functions, such as pairwise hyperbolic distances between distant points, may still encounter standard precision limits, warranting careful boundary checks in production environments.
- Paper: Poincaré Embeddings for Learning Hierarchical Representations, Maximilian Nickel et al. (2017). Read this foundational treatment of Poincaré-ball embeddings first to understand the model whose numerical capacity and optimization failures the source analyzes.
No sufficiently relevant recommendations were found.
