acceptodds
Under review as a conference paper at ICLR 2027

EnvOS: Auditing Computer-Use Agents When the World Moves On

Abstract

Computer-use agents act on information that can become stale before they commit, but apparent failures under such change can also be produced by the evaluation itself. We introduce EnvOS, a protocol for auditing controlled mid-task state changes: changes fire only after the soon-stale value was served; feasibility, corruptions, isolation and delivery are checked; rules are frozen before rollout; separate code re-derives validity verdicts and the outcome terms supported by archived evidence. We evaluate the finalized protocol prospectively in two workflows under one controlled semantic-browser interface. In clinic booking, Claude Opus 5 succeeded on 0/16 instances with no panel, 2/16 with a cue-and-link panel and 15/16 when the same panel displayed the current location (LIVE–CUE: 13 rescues, 0 harms and 3 ties, block-local Holm p=0.000488; CUE–OFF: block-local Holm p=0.5). At the registered block-local level, displaying the current location increased success relative to this cue-and-link panel. The CUE–OFF contrast did not reject; this does not establish absence of a cue effect. Contradiction, salience, and reduced verification effort remain possible explanations. Over all twelve claim-bearing tests, LIVE–CUE's Holm p=0.00293. In insurance-claim approval, Claude Opus 5 succeeded on 10/10 silent-change instances without review-step re-exposure, Claude Sonnet 5 on 1/10 without and 9/10 with it (8 rescues, 0 harms, 2 ties; registered block-local p=0.00781; retrospective twelve-test Holm p=0.0781). A separately pre-registered follow-up with a non-Anthropic configuration (Gemini 3.1 Pro Preview) did not meet the registered effect criterion (0/10 to 5/10, exact p = 0.0625). Evaluator audits exposed their own limits: on a sealed challenge, validity re-derivation met 24/28 expectations; on a second, specified without implementation or outcome access, the pipeline counted none of 26 corruptions as clean but gave no verdict on 2; on a third, which an external agent designed and generated, it caught more corruptions than raw-log baselines in both workflows. It still counted 12 of 39 as clean. These contrasts estimate within-configuration effects; unrejected comparisons do not establish absence or equivalence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.