acceptodds
Under review as a conference paper at ICLR 2027

OSGuard: A Benchmark for Safety in Computer-Use Agents

Abstract

Computer-use agents can complete benign user instructions while violating important constraints of the user’s environment. We introduce OSGuard, a dual-granularity benchmark suite for evaluating safety through local, pre-execution guardrail decisions and end-to-end task execution. Its action-level benchmark contains 324 human-annotated examples in which guardrails classify candidate actions as allowed, unrelated, or unsafe given the original instruction and current interface state. Its risk-augmented execution suite contains 45 tasks derived from 40 OSWorld tasks, keeping original instructions unchanged while modifying the environment to introduce state-dependent safety constraints and preserve a safe path to completion. Augmented evaluators retain the original task-success criteria and add explicit state-based safety checks, distinguishing safe completion from nominal success that violates these constraints. On the action-level benchmark, the strongest evaluated guardrail reaches 79.9% accuracy and 0.80 macro-F1, but performance drops substantially on actions from risk-augmented executions. In full-task evaluation, an unguarded agent completes 62.2% of tasks safely while 37.8% result in unsafe completion; adding the strongest guardrail reduces unsafe completion to 33.3% while leaving safe success unchanged. These results show that state-dependent safety constraints remain challenging both to recognize locally and to preserve during end-to-end computer use.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.