Slice2Fix: Discovering Error Slices for Targeted GUI Agent Repair
Abstract
Graphical-user-interface (GUI) agents are increasingly used to operate mobile applications, yet they still fail whole groups of tasks that share a common property. Existing ways of learning from execution feedback stay at the level of single tasks or written reports: they repair the tasks that failed, keep the tasks the agent already solves, or summarise the failures into weakness reports, and none of them turns the property behind the failures into new tasks without relying on a stronger model to solve them. To address these limitations, we propose Slice2Fix, which repairs a GUI agent by its systematic failures rather than by its failed tasks. Its core idea is to define a systematic failure as an error slice, a group of tasks described by a combination of controllable task features, so that the object that names a failure also specifies the tasks that repair it, and a reference executor written for each task family solves them, closing the loop from diagnosis to repair. Across three strong open-source end-to-end GUI agents at the 8B scale, slice-directed repair raises success on the repaired slices by 20 to 29 percentage points. With the same amount of data, untargeted tasks from the same generator, applications and operations gain 13.1 points there on GUI-Owl-1.5-8B, against 20.0 for slice-directed repair. The repaired agents also improve on AndroidWorld, including templates that appear in no training data. We also release FeatBench, a feature-labelled benchmark of 974 mobile tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.