acceptodds
Under review as a conference paper at ICLR 2027

Actions Are Not Text: Direct Executable-Set Decoding for GUI Agents

Abstract

GUI agents ultimately execute typed environment commands, yet many policies first serialize each action into text and then parse it back. We ask whether this autoregressive round trip is necessary when the current legal action set is finite and explicit. DirectAct retains the pretrained backbone but replaces action-string generation with a small head that scores complete executable tuples, emitting no action tokens. In an ecological comparison across Qwen2.5-VL-7B, MAI-UI-8B, and Llama-3.1-8B, an unadapted frozen JSON interface emits 70.53–94.78 action tokens and takes 29.70–42.40 times the loaded action-model-path latency of DirectAct. We then test whether this practical gap survives a strong trained comparator by aligning training data, input, parameter scale, and legal support with greedy exact-set LoRA generation. On 1,218 prospective states from 88 officially successful AndroidWorld reference workflows, DirectAct satisfies a predeclared two-percentage-point action-weighted non-inferiority criterion on Qwen and MAI; Llama provides descriptive transfer. Across the three backbones, DirectAct is 4.51–5.13 faster even though the generator emits only 6.20–6.68 tokens. These results identify an interface boundary: language models may construct policy state, while finite environment actions can be selected natively rather than serialized autoregressively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.