acceptodds
Under review as a conference paper at ICLR 2027

SAFEST: Rule-Guided Executable Safety Task Synthesis for Computer Use Agents

Abstract

Computer-Use Agents (CUAs) extend large language models with the ability to perceive screens, manipulate files, and operate multiple applications on behalf of the user. While modern LLMs receive extensive chat-style safety alignment, we demonstrate that this alignment does not transfer reliably to the CUA setting: a model that declines harmful textual queries often executes the same intent when instantiated as a concrete computer-use task. We address this gap by constructing safety training data tailored to CUA observations, action interfaces, and execution environments. To this end, we introduce , a afety-ware ramework for xecutable tructured raining data synthesis, which employs a hierarchical pipeline to progressively ground abstract risk categories into concrete, application-specific adversarial tasks and validated safe-execution trajectories. Leveraging this framework, we construct a safety fine-tuning dataset comprising 2,724 executable tasks (1,299 direct and 1,425 indirect) that span 8 core applications and cover 43 distinct harmful categories. Fine-tuning two open-source CUA backbones (Qwen3-VL-8B and EvoCUA-8B) on this dataset significantly enhances their safety alignment: the average safe rate across four CUA benchmarks rises from roughly 41% to 87.8–90.1%, while over-refusal on benign tasks remains low and accuracy on general GUI benchmarks is preserved. Our code and dataset are available at https://anonymous.4open.science/r/SAFEST-42F2.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.