acceptodds
Under review as a conference paper at ICLR 2027

ELDER: LEARNING COMPACT EVIDENCE FOR LONG- DOCUMENT VISUAL REASONING

Abstract

Long-document visual reasoning is constrained by visual-token budgets, as processing full pages allocates substantial input capacity to content unrelated to the question. To address this problem, we introduce Evidence Learning for Document-Element Reasoning (ELDER), a dual-role framework that reduces visual-token input for answer generation while preserving the internal layout of selected evidence. A shared vision-language model selects complete document elements and answers using only their original crops. Supervised fine-tuning uses families of verified evidence sets to initialize both roles, teaching multiple supporting combinations for the same question instead of a single canonical selection. We further formulate Minimal Reliable Evidence Group Relative Policy Optimization (MRE-GRPO) to compare alternative evidence sets by the mean quality of multiple answers generated from each set and its visual-token cost. This set-level comparison can favor complementary combinations and useful corroboration when their answer-quality gains justify the added tokens, and discourage extra content that does not improve sampled answers. A separate answer objective improves reasoning under each fixed set, coupling evidence compression with learning to use retained content. Across five benchmarks, ELDER achieves an average score of 67.6, exceeding the strongest external method with results on all five by 2.3 score points. Compared with answer-only reinforcement learning, ELDER reduces mean answer-input visual tokens by 43.2%, with a 0.3-point decrease in average score.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.