acceptodds
Under review as a conference paper at ICLR 2027

EVA-Reason: Learning Adaptive Evidence Acquisition for Neuroimaging Reasoning

Abstract

Strong multimodal models often struggle to answer difficult neuroimaging questions directly from MRI. We show that these failures are not simply explained by a lack of predictive signal. On the same controlled task, an end-to-end CNN reaches 0.962 balanced accuracy and a Qwen2.5-VL post-projector linear probe reaches 0.897, while direct generation is at chance (0.333). Language-side fine-tuning with frozen vision encoder/projector raises direct generation accuracy to 0.743, narrowing but not closing the representation-to-generation gap. We examine whether explicit quantitative evidence can improve this reasoning bottleneck. We evaluate nine frontier and open-weight multimodal models on the same questions with and without access to quantitative measurement tools. On the hardest questions (L3, the most challenging reasoning level), Tool-OFF accuracy ranges from 36.4% to 56.4%. Four frontier models improve significantly with tools, reaching 75.6% to 80.0% L3 accuracy. Other models improve little or degrade, showing that tool access alone is insufficient. Motivated by these findings, we introduce EVA-Reason, a framework for learning adaptive evidence acquisition in neuroimaging. We construct multi-step evidence-seeking trajectories using Claude Sonnet 4.6 and Gemini 2.5 Pro, and use them to fine-tune a 3B VLM; frontier comparators use the same interface zero-shot. We evaluate EVA-Reason on NeuroQA questions organized into three reasoning levels. On the held-out test set, under normalized exact-match scoring, EVA-Reason achieves 96.4% overall accuracy and 89.3% L3 accuracy. Matched Qwen-3B controls reach 64.5/82.2% overall/L3 without quantitative evidence and 94.6/88.9% when a prespecified relevant-evidence set is supplied upfront; EVA-Reason selects evidence autonomously at inference, without a prespecified per-item evidence set. Under zero-shot exact match, EVA-Reason reaches 89.3% L3 versus 80.0/79.6/79.1% for Sonnet 5/Gemini 2.5 Pro/GPT-5.5; several gaps narrow to non-significance under few-shot prompting or option-prefix scoring. EVA-Reason requires no frontier-model calls at inference and runs locally on a single GPU. These results show that a small open-weight model can learn to actively acquire and use quantitative neuroimaging evidence across multiple decision steps.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.