acceptodds
Under review as a conference paper at ICLR 2027

WISH: A Training-Free World-Model Invocation and Selection Harness for GUI Agents

Abstract

GUI agents can use world models to compare the predicted outcomes of candidate actions, but the evidence they need varies across screens and candidate actions. Existing methods commonly provide the agent with one world model per run and use a fixed pipeline or external rule to decide when to call it. We introduce **WISH**, a training-free harness that lets a frozen GUI agent autonomously decide whether to request world-model predictions and which world model each candidate action needs. To guide these decisions, we abstract recurring evidence needs from teacher explanations into conditions the agent can assess using the current GUI and candidate actions, and combine this guidance with supporting components in WISH. Before requesting a world model's predictions, the agent judges whether those predictions could change its choice among candidate actions. If so, it selects a world model to provide the information each candidate needs. With Qwen3-VL-8B on AndroidWorld, WISH achieves 50.0% task success, exceeding the strongest reproduced baseline by 7.8 percentage points. WISH requests world-model predictions on 21.7% of steps, compared with 23.1% for that baseline, and improves task success over the evaluated baselines across two GUI agents and three benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.