Learning to Reason with Procedural Guidance in MLLMs
Abstract
Reinforcement learning with verifiable rewards (RLVR) can improve reasoning in large language models (LLMs). However, when all sampled responses receive the same reward, the group-relative advantage in methods such as Group Relative Policy Optimization (GRPO) is zero, so the group provides no reward-driven learning signal. Hint-based RL helps the model avoid all-incorrect rollout groups by supplying part of a reasoning trace, usually as free-form text written by an oracle model. Instead of relying on an oracle model, we introduce procedural guidance, in which the program that generates problems and computes their answers also provides correct, controllable guidance. We instantiate this concept in a procedural data generation framework called Vision Reasoning Gym (VRGym), specifically designed for multimodal reasoning problems. The derivations generated by VRGym contain maskable values that can be replaced with an anonymous "?" mark while preserving their reasoning structure. Using these derivations, we design On-Policy Procedurally Graduated Guidance (OPGG), which maintains a scalar guidance-withdrawal budget for each task across training steps and updates it based on each rollout group's accuracy. As the budget increases, OPGG first masks selected values and then truncates the remaining derivation, allowing the model to learn to solve problems without guidance. With this guidance schedule, OPGG achieves the highest seven-dataset macro-average accuracy among the evaluated hint-based methods on both backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.