acceptodds
Under review as a conference paper at ICLR 2027

LatentCUA: A Decoupled World Model for GUI Prediction and Planning

Abstract

Computer-use agents typically choose actions from the current screen without explicitly reasoning about their future effects. We introduce LatentCUA, a compact world model that enables an agent to compare candidate actions through latent prediction. LatentCUA learns a semantic representation of GUI screens, predicts how this representation evolves under an action, and scores imagined future states against the agent's goals. Its modules are trained in separate stages, allowing representation learning, dynamics prediction, and state evaluation to use supervision suited to each role. During planning, a vision-language agent proposes candidate action sequences and LatentCUA ranks their predicted outcomes before execution. Offline ablations validate these design choices, with the final configuration providing the strongest overall balance of representation quality, multi-step prediction, and candidate discrimination. On OSWorld, direct world-model selection improves lower-performing planners, while selective assistance improves all five tested planners. We further provide a decision-sufficiency framework that relates representation, prediction, and scoring errors to action-selection quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.