Hyperbolic Residual Connections for Stability at Scale
Abstract
Large language models (LLMs) have become highly capable at modeling text, largely due to systematic gains from increasing model size and training scale. More recently, LLMs with hyperbolic embedding spaces and operations have shown strong performance by modeling natural language on manifolds that better reflect its latent hierarchical structure. However, hyperbolic models have not been scaled to the same extent as their large pre-trained Euclidean counterparts, in part due to difficulties maintaining training stability as model size increases. In this work, we identify two instabilities of existing hyperbolic residual connections: requiring exceedingly large layer embedding norms to change embedding directions and not being robust to larger layer outputs. We propose a new hyperbolic residual connection, TangentAttnRes, that addresses this instability through two features: (1) aggregation in tangent space, which addresses the first instability by reducing the norm required to make directional changes to the residual vector; and (2) attention-based aggregation, which addresses the second instability by selectively weighting past layer outputs to dampen the impact of unstable layers. We apply our method to multiple hyperbolic LLMs and language Transformers across diverse benchmarks, demonstrating that it improves performance over the existing residual connections. This is especially the case as we scale model parameters, where TangentAttnRes shows continually improved performance while the same scaling causes models with prior methods to show negligible gains or even collapse to near-random performance. Our findings demonstrate that TangentAttnRes is crucial for achieving further performance gains as hyperbolic LLMs scale, presenting a path toward models that can compete at the frontier.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.