DERIVE: Dependency-Aware and Evidence-Responsive Reasoning via Intervention and Verifiable Execution for MLLM Long-Document Understanding
Abstract
Multimodal large language models (MLLMs) typically tackle long-document reasoning by retrieving sparse evidence from abundant irrelevant pages before reasoning over it. Yet our analysis shows that substantial errors persist even when gold evidence is directly provided, shifting the challenge from locating evidence to reasoning reliably from it. However, existing methods mainly verify final outcomes or individual reasoning states, which does not establish whether downstream reasoning actually depends on the answer-determining evidence. We therefore propose Dependency-aware Evidence-responsive Reasoning via Intervention and Verifiable Execution (DERIVE), which makes evidence-to-reasoning dependencies explicit and executable through Verifiable Execution (VE) chains, and makes them directly testable through controlled Evidence Intervention (EI). VE chains expose evidence provenance and intermediate execution to deterministic verification, while EI modifies answer-determining evidence and re-executes the affected downstream reasoning to produce verifiable dependency responses. DERIVE is trained with VE-chain SFT establishing the grounded executable reasoning format, followed by GRPO with executable verification over original and intervened evidence states to reinforce grounded, executable, and evidence-responsive reasoning. Extensive experiments on six public benchmarks demonstrate that DERIVE consistently outperforms competitive baselines, especially on arithmetic-intensive tasks requiring evidence-dependent reasoning. Code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.