acceptodds
Under review as a conference paper at ICLR 2027

Gold-LP: Data Selection for Vision-Language Reinforcement Learning with Verifiable Rewards

Abstract

Reinforcement learning with verifiable rewards improves vision-language models, and which questions enter training affects what the model learns. Choosing them is itself costly: selectors that estimate each candidate's difficulty from sampled responses must generate over the whole pool before optimization begins, a fifth as many responses as the training run itself, paid whether or not the resulting ranking helps. Existing selectors fall into three categories: sampling responses; reusing an earlier training run's outcomes, which requires one to exist; and substituting the model's own probability, which reduces generation by an amount the published papers do not always make clear. However, these approaches leave a fundamental question unanswered: can a useful subset be identified without generating a single response, and how much accuracy can such a subset deliver? We introduce Gold-LP: one forward pass per candidate scores the initial model's likelihood for the correct option letter under a fixed answer prefix, and a contiguous interval of the resulting ranking is retained, leaving group relative policy optimization unchanged. Training Qwen3-VL-8B-Instruct on 1,000 of 3,023 DOCCI-derived examples at a matched step budget, it reaches the highest mean accuracy among the evaluated fixed-size selections on both holdouts across three runs; the geometry gain over random selection is 1.35 points (unadjusted 95% CI [0.41, 2.30]). Which region is kept matters as much as how candidates are scored: with score and subset size held fixed, accuracy peaks at our interval rather than at the lowest-score tail on both holdouts, exceeding a published mean-centred rule in distribution by roughly four times the across-seed deviation, while three fixed-size comparators adapting published selectors do not reach it. Selection-time generation can therefore be removed entirely without giving up downstream accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.