acceptodds
Under review as a conference paper at ICLR 2027

TraceGain: Marginal Answer Utility for Document VLM Reveal-or-Stop Control

Abstract

At a visited document state, stopping and acquiring another high-resolution region are mutually exclusive actions, yet region saliency and answer confidence score them in incompatible units. TraceGain turns this choice into a measurable control problem by assigning every legal zoom and shift its realized one-step change in official answer score and placing stop at the exact zero of that scale. Exhaustive counterfactual expansion of student-visited states supplies the targets; a centered dense teacher shapes only the offline regression target, and a 1.35M-parameter head batches all legal actions at inference. On DocVQA, InfographicVQA, and ChartQA, raw predictions attain Spearman –, Utility ECE –, and one-step stop AUPRC –. Against DirectPolicy matched for inputs, capacity, teacher privilege, roll-ins, and budget, fixed-cache point estimates show 57% more stop errors and 65% more regret for direct classification. Matched endpoint gains span 0.3–1.1 points with LLaVA and 0.3–1.0 with Qwen2.5-VL; in the DocVQA operating study, TraceGain reaches 81.7 ANLS at 968 reported pre-encoding tokens and 438 ms, compared with 82.4 at 2,016 tokens and 847 ms for Full-Res. A common marginal-utility coordinate thus connects action selection, stopping, calibration, and compute allocation for short-horizon document inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.