ELDER: LEARNING COMPACT EVIDENCE FOR LONG- DOCUMENT VISUAL REASONING
Abstract
Long-document visual reasoning is constrained by visual-token budgets, as processing full pages allocates substantial input capacity to content unrelated to the question. To address this problem, we introduce Evidence Learning for Document-Element Reasoning (ELDER), a dual-role framework that reduces visual-token input for answer generation while preserving the internal layout of selected evidence. A shared vision-language model selects complete document elements and answers using only their original crops. Supervised fine-tuning uses families of verified evidence sets to initialize both roles, teaching multiple supporting combinations for the same question instead of a single canonical selection. We further formulate Minimal Reliable Evidence Group Relative Policy Optimization (MRE-GRPO) to compare alternative evidence sets by the mean quality of multiple answers generated from each set and its visual-token cost. This set-level comparison can favor complementary combinations and useful corroboration when their answer-quality gains justify the added tokens, and discourage extra content that does not improve sampled answers. A separate answer objective improves reasoning under each fixed set, coupling evidence compression with learning to use retained content. Across five benchmarks, ELDER achieves an average score of 67.6, exceeding the strongest external method with results on all five by 2.3 score points. Compared with answer-only reinforcement learning, ELDER reduces mean answer-input visual tokens by 43.2%, with a 0.3-point decrease in average score.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.