Reward Reasoning for Its Gain over a Snap Judgment: Inverse-RL Cold Start of GUI Agents with an Adversarial Critic
Abstract
Cold-starting GUI agents requires learning from expert demonstrations before sparse task-completion rewards can support effective online reinforcement learning. Supervised fine-tuning (SFT) provides this initialization, but fitting expert traces can come at the cost of general capability. Learning a reward from demonstrations offers another route: the agent trains on its own decisions, guided by a critic that reasons about whether a decision came from the expert or the policy. The challenge is that source labels supervise the critic's final answer, not whether its reasoning helped. Obvious errors and stylistic cues can make a pair easy to classify, allowing uninformative reasoning to receive the same reward as reasoning that improves the judgment. We introduce ReGain, which trains the critic on the improvement its reasoning provides over a direct, no-reasoning judgment. For each policy-expert pair, we score the verdict distributions with and without chain-of-thought reasoning using a detached Brier score and reward their difference. The continuous score captures improvements even when the selected verdict is unchanged, while centring rewards across distinct pairs gives greater relative credit to reasoning that improves the comparison. Confidently correct direct judgments leave little room for positive gain, reducing the reward for reasoning that merely repeats obvious source cues. On a 27B GUI agent trained on 20,480 demonstrations, ReGain outperforms the adversarial baseline RARO and SFT on the same data and delays the critic's reliance on stylistic cues. It matches SFT trained on nine times as many demonstrations in pass@1 and pass@4, without the 14% loss of general capability incurred by either SFT baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.