acceptodds
Under review as a conference paper at ICLR 2027

SEEING IS NOT RELATING: LEARNING RELATIONAL REPRESENTATIONS FOR MULTI-IMAGE UNDERSTANDING

Abstract

Multi-image understanding requires vision–language models to connect visual information across images. Controlled diagnostics across four model families show that models can correctly recognize task-relevant local details, including attributes and within-image relations, yet still fail at cross-image comparison. Complementary information interventions show that providing the correct relation target substantially improves task performance, highlighting the importance of reliable cross-image relations. Motivated by these findings, we construct structured supervision covering entities, attributes, within-image predicates, and cross-image comparisons to learn relational features that complement pretrained visual content. We propose Relation-Aware Dual-Stream Fusion (RDSF), which learns a dedicated relation stream, exchanges context across images, and integrates the resulting features with native visual tokens through token-aligned residual fusion. The pretrained backbone remains frozen throughout training. Across four Qwen3.5 scales, RDSF improves the six-benchmark mean by 1.99–6.51 percentage points. Additional evaluations on InternVL3.5 and LLaVA-OV1.5 show gains in the four-benchmark mean, supporting the applicability of complementary relational learning across model architectures.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.