acceptodds
Under review as a conference paper at ICLR 2027

SafeSym: Generating a Safety-Constrained Symbolic World Model for GUI Agents

Abstract

GUI agents can execute complex multi-step tasks in real-world digital environments, but ensuring their safety remains a key challenge. Existing approaches rely on post-hoc guardrails, which fail to enforce safety during decision-making. Symbolic world models provide a structured representation of environment dynamics and enable planning with explicit reasoning, making them a promising foundation for controllable and interpretable agent behavior. However, existing symbolic approaches do not explicitly model safety, leading to potential risks during planning and execution. To address this, we propose SafeSym, a novel multi-agent framework that constructs safety-constrained symbolic world models and embeds safety constraints directly into planning. SafeSym builds and verifies symbolic models represented in the Planning Domain Definition Language (PDDL), dynamically instantiates task-specific constraints via rule matching, and performs planning over the constrained model to ensure safety by construction. Extensive experiments on GUI-based tasks show that SafeSym achieves the highest overall safety compliance and success-with-obligation rate, outperforming existing SOTA methods while maintaining competitive task success and requiring zero runtime safety guard calls. These results demonstrate that integrating safety into symbolic planning provides an effective and interpretable solution for safe GUI agents. The source code is available at https://anonymous.4open.science/r/SafeSym.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.