acceptodds
Under review as a conference paper at ICLR 2027

VisCal: Calibrating Visual Intermediate-State Capacity by Downstream Task Utility

Abstract

Interleaved visual reasoning models generate intermediate visual states. Yet the number of visual tokens or latents assigned to each state is usually fixed or calibrated by reconstruction fidelity. Holding reasoning trajectories fixed while varying capacity, we find that fidelity does not reliably order downstream answer utility, useful capacity depends on the problem and reasoning state, and capacity effects interact across steps. Visual capacity is therefore a budgeted sequential decision rather than a representation-quality parameter. We introduce VisCal at two levels: VisCal-Ext keeps the backbone frozen and distills hindsight trajectory allocations into a lightweight causal controller, isolating the deployment value of utility calibration and joint allocation. VisCal-OPD places a budget-conditioned capacity action before visual generation and constructs an outcome-calibrated teacher from the correct answer and complete outcomes of structured cross-step branches. Distilling this teacher jointly into capacity, visual, and text policies lets the model learn what each visual state contains and how later reasoning uses it. Across four models and eight transfer benchmarks disjoint from the post-training pool, VisCal-OPD improves matched-budget accuracy by 5.14 percentage points over the released models and matches their accuracy with an estimated 65.1% fewer visual tokens or latents from one budget-conditioned checkpoint per model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.