acceptodds
Under review as a conference paper at ICLR 2027

Sampling Versus Longer Responses for Visual Reasoning: A Matched-Prompt MathVista Study

Abstract

Visual reasoning can receive additional inference budget through longer responses or repeated sampling, but these interventions need not have the same effect. We compare a 512-token single response, a 2,048-token single response, and voting over four 512-token responses on the same 300 multiple-choice MathVista questions, using Qwen2.5-VL-3B and 7B with matched prompts and image preprocessing. Under the fixed explicit-answer scorer, four-response voting improves accuracy by 15.3 and 6.7 percentage points, respectively; paired question-bootstrap 95% intervals are [11.3, 19.7] and [3.7, 10.0]. Increasing the single-response ceiling yields gains of 2.7 and 2.0 points with intervals spanning zero. Actual generated-token counts explain why equal ceilings are not equal compute: four-response voting uses roughly four times as many output tokens, whereas most single responses finish well below either ceiling. A broader, label-blind answer-parser sensitivity and an average-single-path baseline preserve the voting gains. Most newly correct voting outcomes arise on questions where the first response lacks a parsed final answer. A separate analysis of larger checkpoints uses saved predictions; incomplete response records prevent uniform regrading. The findings support repeated sampling for these two checkpoints and this subset; they establish neither a general model-size scaling law nor a compute-efficiency advantage at equal realized cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.