acceptodds
Under review as a conference paper at ICLR 2027

Omni-Modal LLMs Know More Than They Say

Abstract

Omni-modal large language models (Omni-LLMs) extend LLMs to process heterogeneous modalities (*e.g.*, vision, audio, and IMU signals) alongside language, enabling a broader range of multimodal tasks. Yet we find that these models often fail to express information from non-textual modalities even when they perceive it correctly. We term this disconnect *knows-but-cannot-say*. To investigate this phenomenon, we introduce Token-level Modality Alignment (TMA), which measures the alignment between individual modality representations and response representations. Our analysis shows that non-textual representations are less aligned with response tokens and receive less attention than textual representations, suggesting a potential explanation for uneven modality preferences. Motivated by these findings, we propose Omni-Catalyst, a simple *perceive* *infer* strategy that complements the original modality inputs with self-generated text to make non-textual knowledge more accessible. Across seven Omni-LLMs, Omni-Catalyst consistently improves performance (+7.29%) and yields more balanced modality preferences, with further gains from diverse training strategies. Further analysis demonstrates that the knowledge extraction phases increase TMA, suggesting that improved alignment makes non-textual information more accessible during response generation. In summary, our findings uncover **how Omni-LLMs can better express modality knowledge through language**.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.