CAVEAT: Two-View Auditing with a Deterministic Veto Lowers GUI-Agent Verifier False Positives Below Recorded LLM Judges
Abstract
Models that read a recorded GUI-agent trajectory and predict task success accept narration as evidence. Most of a record audit’s false-positive reduction comes from a deterministic structural check, not from a second model view: of the expert-false test rows the view that reads the agent’s answer reports 67.7%, both views 58.1%, and the check 22.6% (78.6% of the drop), at 50.0% recall against the best recorded judge’s 93.5%. We present CAVEAT (Counterfactual Agent Verification via Evidence-only Auditing of Trajectories), the selective verifier that combines them: it reports success only when both reasoning-free views agree and the deterministic checker finds no structural contradiction. On the untouched test split the audit rule reports at 42.7% coverage with 81.6% precision (22.6% FPR), while the strongest measured LLM-judge, critic, reward-model, or single-view baseline reaches 73.4% precision and none gets below 35.5% FPR; the predeclared 10% risk certificate fails calibration (95% upper bound 65.4%), so the deployed policy abstains rather than overclaim. Because neither view reads agent reasoning, rewrites confined to reasoning cannot change a verdict; we prove this but do not measure it, since every implemented gaming attack also rewrites the terminal answer, and those attacks raise the gate’s false-positive rate on the pilot from 22.2% to 66.7%. The certificate fails partly by construction: the target needs 29 task-level reports and the rule produced 27. A live layer ran step ablation and state mutation on restored WebArena/VisualWebArena snapshots for a 30-trajectory design pilot (ablation returned a verdict on 10 of 22 replayable trajectories, mutation on none); an ablation-confirmed claim was expert-false, so we do not equate a verification record with truth. A 7B LoRA student reproduces 83.3%–88.9% of the frozen auditor’s view verdicts (3 seeds) at 6.04–6.42× its latency and fails its own certificate at every seed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.