Lorentz Busemann Language Models for Efficient Hyperbolic Language Modeling
Abstract
Prior studies suggest that token level structure is hierarchical and non-Euclidean, motivating hyperbolic large language models (LLMs). Fully hyperbolic models such as HELM add hyperbolic operations throughout the Transformer, increasing training and inference costs. Besides, evidence for hyperbolic structure mainly concerns token embeddings and last hidden states, which together determine the logits at the output layer. We therefore propose the Lorentz Busemann Language Model (LBLM), a hyperbolic LLM that retains standard Transformer backbone and introduces learnable hyperbolic token embeddings at the output layer. LBLM uses the Lorentz Busemann Logit (LBL) to score tokens with the Busemann function, treating hidden states as boundary directions and avoiding their exponential mapping into hyperbolic space. To optimize the token embeddings under learnable curvature, we develop Curvature Learning Riemannian Adam (CRAdam), which keeps the embeddings in the hyperbolic space and preserves the moment norm across curvature changes, without repeated logarithmic and exponential maps through the tangent space. Across scales from 0.3B to 1.4B parameters, LBLM improves mean accuracy by up to 7.8% relative to Euclidean baseline trained with the same recipe, with at most 5.0% training overhead. Compared with the fully hyperbolic HELM, which cannot be trained stably under this recipe, LBLM attains higher accuracy with each model trained under its own recipe. Under identical workloads, LBLM achieves at least 5.3 the training throughput and 2.8 the inference throughput of HELM. LBLM offers a simple and efficient hyperbolic LLM, showing that hyperbolic space need not permeate the Transformer backbone to benefit language modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.