AsyPO: Asymmetric Post-Training for GUI Agents with Segment-Level Reasoning Guidance
Abstract
Graphical user interface (GUI) agents require both executable actions and reasoning, but existing outcome-based post-training provides only coarse supervision for the latter. Existing post-training methods mainly include reinforcement learning (RL) and on-policy self-distillation (OPSD). Although OPSD enables process-level supervision through teacher–student token probabilities, applying it independently to long reasoning spans can distort their global structure. We propose an asymmetric post-training framework that combines action-level Group Relative Policy Optimization with relative segment-level reasoning guidance. Our method aggregates teacher guidance within reasoning segments, centers the resulting signals across each reasoning span, and selectively applies them through negative-advantage gating. This provides targeted reasoning supervision while preserving the overall structure of model-generated reasoning. Experiments on OSWorld and OmniACT demonstrate consistent improvements over the evaluated baselines at both 4B and 8B model scales.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.