acceptodds
Under review as a conference paper at ICLR 2027

Prompt Optimization as a Probe for Data-Efficient Multimodal GRPO

Abstract

Group Relative Policy Optimization (GRPO) can substantially improve multimodal reasoning, but its reliance on multiple rollouts per training problem makes data-efficient training important. Existing data-selection methods typically estimate sample utility from signals observed under a fixed prompt, providing only one view of the current policy's behavior. We introduce PACE (Prompt-Induced Answer Change Estimation), which repurposes prompt optimization as a controlled behavioral probe for GRPO data selection. PACE keeps the initial policy fixed and compares answer correctness under the original and optimized prompts. The resulting bidirectional transitions—unlocks from incorrect to correct and regressions from correct to incorrect—define a prompt-sensitive boundary that captures how the policy responds across elicitation conditions. PACE aggregates these transitions within semantic clusters, retains regions where optimized prompting has a positive aggregate effect, and applies balanced selection to construct a compact training subset. The optimized prompt is used only for probing and selection; downstream GRPO training and evaluation use the original prompt. Experiments using MM-Eureka for training and seven multimodal mathematical reasoning benchmarks for evaluation show that PACE selects only 5,000 examples, approximately 9.1% of the full training set. On Qwen2.5-VL-7B-Instruct, PACE achieves an average accuracy of 51.1, compared with 50.7 for Full-Data GRPO trained on 54,931 examples, while outperforming fixed-budget, larger-budget, and online selection baselines. PACE remains effective across Qwen2.5-VL-3B and InternVL3-2B. Further analyses show that the prompt-sensitive boundary captures information complementary to conventional fixed-prompt signals, that both transition directions provide useful training signals, and that prompt-sensitive filtering improves selection strategies. These results demonstrate that prompt-induced behavioral responsiveness provides an effective signal for data-efficient multimodal GRPO.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.