acceptodds
Under review as a conference paper at ICLR 2027

Speculative Rollback Correction: Counterfactual Credit Assignment for GUI Agent Imitation

Abstract

Interactive imitation learning supervises agents on learner-induced states, yet conventional methods typically identify errors via disagreement with a reference policy.In multi-solution GUI tasks, this proxy frequently conflates genuinely harmful actions with valid alternative trajectories. To resolve this ambiguity, we define errors through task recoverability: an action is harmful to the extent that it reduces the maximum probability of eventual success achievable by any continuation policy. Because recoverability optimizes over policy space, it reflects an intrinsic property of the task and environment rather than the choices of a specific teacher. We show that the resulting recoverability process admits an exact Doob decomposition, whose predictable compensator yields per-step harm and provides a principled mechanism for localizing feasibility-reducing actions. We instantiate this theoretical framework as Speculative Rollback Correction (SRC). In SRC, the student agent executes a short speculative trajectory segment, after which the teacher localizes the earliest action causing a drop in recoverability. A reset-and-replay mechanism then restores the corresponding state for targeted correction, followed by a hard verifier that filters candidate continuations before updating a standard next-action policy. Across WebArena-Infinity, WebArena-Lite, and OSWorld, augmenting Expert SFT with SRC data improves teacher-free deployment success rates by 9.7, 3.6, and 12.9 percentage points, respectively. Under comparable teacher-query budgets, SRC outperforms OEC-style training across web benchmarks; compared to LEAP-style training, it achieves superior success on WebArena-Infinity and competitive performance on WebArena-Lite (within 2.6 points) while reducing teacher queries by 56–58%. Controlled experiments further confirm the efficacy of corrective labels and teacher-guided localization. At test time, the deployed policy operates autonomously without requiring teacher queries, verifiers, or environment resets, establishing recoverability-guided correction as a practical data-collection principle for resettable GUI environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.