GEAR: Co-evolving Grounded VLM Agents with Environment-Aligned Rubrics
Abstract
Training multi-turn VLM agents from environment feedback remains challenging because outcome rewards indicate whether a trajectory succeeds, but provide limited supervision for why intermediate visual reasoning steps and actions succeed or fail. Existing dense process rewards are often manually specified or fixed, which makes them difficult to adapt to the evolving trajectory distribution induced by RL. We propose GEAR (Grounded Environment-Aligned Rubrics), a co-evolutionary training framework that jointly optimizes a VLM policy and a rubricator that self-generates grounded rubrics for spatial grounding, action legality, temporal consistency, and transition progress. Each rubric is used as an auxiliary process reward only when its satisfaction is statistically aligned with grounded environment feedback, such as task success, return, legality, or progress, which prevents self-generated supervision from drifting toward textual self-consistency. Experiments on multi-turn VLM agent benchmarks show that GEAR improves over SFT and outcome-centric RL baselines, with ablations confirming that environment-aligned rubric evolution is important for stable and interpretable policy improvement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.