LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI Agents
Abstract
Modern GUI-agent frameworks achieve strong desktop task performance with frontier API models, yet persistent control information often remains implicit in growing interaction trajectories. At each step, the planner reconstructs the active task stage, accumulated evidence, and runtime feedback before deciding the next action. This dependence becomes more pronounced under weaker local reasoning backbones. Across four representative state-of-the-art frameworks, replacing GPT-5 with Qwen3.5-9B reduces average OSWorld SR-100 from 60.9% to 35.2%. Trajectory annotation further identifies at least one control failure in 94.7% of failed trajectories. To address this problem, we introduce LocalLSTC, a training-free architecture that organizes control by temporal scope, maintaining persistent cross-step state to guide short-term execution commitments. The framework comprises two complementary mechanisms. Long-to-Short Planning forms each commitment from persistent state, while Short-to-Long Control integrates execution outcomes back into that state for progress assessment, recovery, and termination. With Qwen3.8-27B, LocalLSTC reaches 73.7% SR-100 on OSWorld and 68.5% on WindowsAgentArena, achieving the strongest local OSWorld result and a new state of the art on WindowsAgentArena. Ablations further support contributions from both mechanisms. Together, these results show that temporal organization of control information can reduce cross-step control failures and complement backbone scaling in locally deployed GUI agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.