acceptodds
Under review as a conference paper at ICLR 2027

GEAR: Co-evolving Grounded VLM Agents with Environment-Aligned Rubrics

Abstract

Training multi-turn VLM agents from environment feedback remains challenging because outcome rewards indicate whether a trajectory succeeds, but provide limited supervision for why intermediate visual reasoning steps and actions succeed or fail. Existing dense process rewards are often manually specified or fixed, which makes them difficult to adapt to the evolving trajectory distribution induced by RL. We propose GEAR (Grounded Environment-Aligned Rubrics), a co-evolutionary training framework that jointly optimizes a VLM policy and a rubricator that self-generates grounded rubrics for spatial grounding, action legality, temporal consistency, and transition progress. Each rubric is used as an auxiliary process reward only when its satisfaction is statistically aligned with grounded environment feedback, such as task success, return, legality, or progress, which prevents self-generated supervision from drifting toward textual self-consistency. Experiments on multi-turn VLM agent benchmarks show that GEAR improves over SFT and outcome-centric RL baselines, with ablations confirming that environment-aligned rubric evolution is important for stable and interpretable policy improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.