Gradient Eligibility Is Not Learning Utility: Hidden Curricula in Vision–Language RL
Abstract
General-purpose vision–language reinforcement learning often trains a single policy across heterogeneous perception and reasoning tasks using group-relative advantages with binary correctness feedback. Only rollout groups containing both correct and incorrect answers yield nonzero relative advantages through this outcome channel, making each query's contribution policy-dependent. Dynamic sampling amplifies this selection by discarding zero-variance groups and refilling the batch, so queries that remain mixed accumulate more updates, creating a hidden curriculum. We ask whether this persistence signals useful learning opportunities or merely reflects continued outcome variability. Before RL, we predict which queries will remain mixed using initial solvability and response demand, measured as the thinking length elicited by a fixed probe VLM. Across Qwen2.5-VL and Qwen3-VL, higher demand predicts longer persistence after accounting for initial solvability and query composition. Under dynamic sampling, the highest-demand decile consequently passes the filter – as often as the lowest-demand decile. Persistence, however, does not imply learning utility. Across matched three-seed interventions at two policy states, high-demand pools pass the filter more often and consume more generated response tokens. Starting from the base policy, both arms improve, but low-demand training gains percentage points more at pass@; all six paired pass@ comparisons favor low-demand training. A past-only inverse-acceptance controller reduces exposure drift by at near-matched cost, with higher observed pass@ and pass@. Thus, controlling which queries receive updates is distinct from deciding which updates are useful.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.