acceptodds
Under review as a conference paper at ICLR 2027

Knowing Where Is Not Knowing What: Evidence-Grounding Failure in Social Gaze

Abstract

Understanding social gaze is essential for socially interactive agents, yet a look's function depends less on its direction than on the subsequent interaction. While multimodal large language models (MLLMs) are evaluated on gaze following, whether they ground functional judgments in this interaction remains largely unexplored. To address this gap, we present **EviGaze**, a benchmark specifically designed to diagnose evidence grounding in social gaze understanding. EviGaze annotates 5,120 gaze events from four social-video corpora with their *function* and *evidence interval*. Since this interval follows the gaze, removing it withdraws the grounds for judgment while keeping the look visible. We define *evidence-grounding failure* (EGF) as label retention after removal once the target and evidence are located. Extensive experiments are conducted to benchmark fifteen MLLMs on EviGaze. Experimental results reveal that EGF persists across all fifteen MLLMs, with evidence removal lowering model accuracy by only 16.0 percentage points against 64.4 for humans. Probing reveals that model states encode functional information that judgments underuse, a dissociation consistent across all seven analyzed models. Prompting models to describe the evidence before answering reduces the EGF rate beyond a matched placebo. The benchmark and code will be publicly released to support further research.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.