acceptodds
Under review as a conference paper at ICLR 2027

Learning to Zoom at the Right Difficulty: Dynamic Curricula for Visual Tool-Use Reinforcement Learning

Abstract

Crop-and-zoom tools enable vision-language models to acquire fine-grained visual evidence and answer questions that require visual detail beyond their native input resolution. Prior work shows that reinforcement learning on tasks solvable without visual tools can induce overfitting rather than improve tool use capability, motivating training on hard tool-essential data. However, training on challenging datasets can be highly inefficient under Group Relative Policy Optimization (GRPO), as filtering out numerous uninformative samples with all-zero rewards requires substantial additional rollout generation. This overhead is especially high when visual tool use involves multiple rounds of generation, tool execution, and visual feedback. We propose Calibrated Dynamic Difficulty Shaping (CalDS), an adaptive sampling framework that makes efficient use of hard examples by matching training-batch difficulty to the model's evolving capabilities. CalDS organizes examples into dynamically updated buckets based on tool-augmented accuracy. It uses a bounded reward-feedback controller with difficulty shaping to adjust sampling probabilities based on bucket difficulty and reward history, emphasizing challenging but currently learnable examples. To account for stale difficulty estimates as the policy evolves, CalDS reuses recent rollouts to calibrate bucket difficulty estimation. On VisualProbe, a challenging high-resolution VQA benchmark, CalDS improves both answer accuracy and crop localization accuracy at a much lower training cost. These results highlight adaptive difficulty control as a means of learning effective visual tool use efficiently from hard training data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.