The Broken Geometry of Hyperbolic Vision-Language Models
Abstract
Hyperbolic vision-language models (VLMs) leverage hyperbolic geometry to model hierarchical structure in multimodal representations. Recent works contrastively align images and text in a shared hyperbolic space while employing entailment cones to impose hierarchical orderings among whole and part representations. However, we find that existing hyperbolic VLMs operate in a near-Euclidean regime, with embeddings concentrated near the origin on manifolds with small curvature. Thus, most embeddings fall outside the theoretically valid domain of entailment cones, suggesting that the intended geometry and entailment mechanism are not fully exploited in practice. Motivated by our observation, we propose a simple Euclidean-based hierarchy-aware VLM that retains the same compositional supervision while eliminating hyperbolic operations and entailment cones. Despite its simpler geometry, our model consistently outperforms existing hyperbolic VLMs across diverse downstream tasks. The empirical results suggest that the gains reported by prior studies may stem primarily from additional compositional data and hierarchical supervision rather than from hyperbolic geometry itself, highlighting the need to carefully examine the contribution of hyperbolic geometry in hierarchical representation learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.