Considering the Forest and Every Leaf: Various Interleaved Visual Inputs for Abstractive Analysis
Abstract
Abstractive visual reasoning (AVR) requires multi-modal large language models (MLLMs) to interpret human-defined visual representations, such as mind maps and relational diagrams. Existing AVR benchmarks primarily study individual abstract images, leaving open how models connect concrete visual evidence with abstract relational structures distributed across multiple images. We introduce Viviana, a benchmark for reasoning over sequences composed of entity images, visualized knowledge triples, and subgraph diagrams. Viviana contains 22,568 instances across 14 tasks, ranging from visual grounding to cross-image relational and graph reasoning. We also construct local-to-global chain-of-thought supervision and propose Gated Knowledge-informed GRPO (GKGRPO), which combines answer rewards with local and global knowledge rewards through an answer-conditioned, two-level gate. On Qwen3-VL models, GKGRPO achieves the highest overall score among the evaluated post-training methods. The resulting models also improve on established multi-image understanding and abstract visual reasoning benchmarks, indicating that structured cross-image supervision can transfer beyond the visual knowledge graph setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.