acceptodds
Under review as a conference paper at ICLR 2027

AVTextBench: Benchmarking Cross-modal On-Screen Text Grounding in Audio-Visual Large Language Models

Abstract

On-screen text has become a widely applied tool to convey video information more effectively, such as subtitles. Despite the success of audio-visual large language models (LLMs), it remains unclear whether they can ground on-screen text in the surrounding audio-visual context. In this work, we show that audio-visual LLMs perform poorly in understanding the relationships between on-screen text and audio-visual streams. To systematically study this problem, we introduce a benchmark for probing these capabilities. Specifically, we introduce a controlled framework covering three fundamental relationships: corresponding, complementary, and independent. This framework provides a comprehensive evaluation of whether audio-visual LLMs can correctly match, integrate, and segregate on-screen textual and audio-visual information in real-world videos. We evaluate our benchmark on state-of-the-art audio-visual LLMs, including both open-source models and commercial models. Experimental results show that existing open-source audio-visual LLMs fail to ground on-screen text in context, with performance near random guessing in multiple tasks. On the other hand, the Gemini-3.8-Flash model achieves strong accuracy, highlighting a large gap between open-source and commercial models. Our work uncovers an unexplored limitation in current audio-visual LLMs and provides new insights for future improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.