VERGE: Verified Early-exit Reasoning with Grounded Evidence in Medical Multimodal Reasoning
Abstract
Long medical reasoning traces can contain a correct intermediate prediction followed by redundant or harmful continuation. We introduce VERGE (Verified Early-exit Reasoning with Grounded Evidence), a GRPO-based framework that uses intermediate verification to supervise credit assignment and reasoning termination. A reasoning continuation probe forces candidate prefixes to transition to answering and checks the resulting predictions against task-specific verification criteria. Accepted boundaries guide advantage redistribution over the original rollouts and supervise the reasoning-termination token; probe completions are used only for verification. VERGE handles discrete answer correctness and continuous localization quality, preserving each grounding rollout's mean GRPO advantage. Starting from Qwen3-VL-4B-Thinking, VERGE improves average accuracy across four medical QA benchmarks from 54.17% to 57.54%. Under a matched 4,096-token protocol, it outperforms Vanilla GRPO by 4.95 percentage points in MedMCQA accuracy and 5.61 points in grounding mIoU on the evaluated subset, while reducing mean output tokens by 77.2% and 69.6%, respectively. These results support learning to terminate reasoning from verified intermediate predictions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.