acceptodds
Under review as a conference paper at ICLR 2027

Seeing What Changed: Transition-Guided Credit Assignment for GUI Agents

Abstract

Reinforcement learning with verifiable rewards (RLVR) for multi-turn GUI agents provides terminal outcome supervision but leaves credit assignment across individual turns unresolved. Uniformly broadcasting a trajectory-level advantage can reinforce both useful actions and mistakes that the agent subsequently corrects. Screenshots before and after each action expose state changes that help distinguish progress from detours, even within successful trajectories. We propose Observable-Transition Advantage Decomposition (OTAD), which uses this evidence to guide turn-level credit assignment. A vision-language model provides hindsight ordinal judgments of each transition conditioned on the verified outcome. OTAD maps these judgments to non-negative weights on the trajectory-level advantage rather than additional rewards: transition evidence sets the magnitude of each turn's advantage while the terminal verifier fixes its sign. This design removes the need for a learned value model and for extra environment interaction. The same judgments guide budget-constrained turn selection and prioritize tasks with discriminative transition structure, improving training-data efficiency. On OSWorld-Verified, OTAD raises success from 42.2% under budget-matched GRPO to 52.1%; at 8B parameters, it surpasses ComputerRL-9B and OpenCUA-72B, while also improving GUI grounding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.