acceptodds
Under review as a conference paper at ICLR 2027

Considering the Forest and Every Leaf: Various Interleaved Visual Inputs for Abstractive Analysis

Abstract

Abstractive visual reasoning (AVR) requires multi-modal large language models (MLLMs) to interpret human-defined visual representations, such as mind maps and relational diagrams. Existing AVR benchmarks primarily study individual abstract images, leaving open how models connect concrete visual evidence with abstract relational structures distributed across multiple images. We introduce Viviana, a benchmark for reasoning over sequences composed of entity images, visualized knowledge triples, and subgraph diagrams. Viviana contains 22,568 instances across 14 tasks, ranging from visual grounding to cross-image relational and graph reasoning. We also construct local-to-global chain-of-thought supervision and propose Gated Knowledge-informed GRPO (GKGRPO), which combines answer rewards with local and global knowledge rewards through an answer-conditioned, two-level gate. On Qwen3-VL models, GKGRPO achieves the highest overall score among the evaluated post-training methods. The resulting models also improve on established multi-image understanding and abstract visual reasoning benchmarks, indicating that structured cross-image supervision can transfer beyond the visual knowledge graph setting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.