acceptodds
Under review as a conference paper at ICLR 2027

Evidence-First Agentic Verification

Abstract

Long-horizon, semi-verifiable agentic tasks often rely on rubric-based evaluation, but the evidence required to assess success can be distributed across the environment and execution trace. When rubric verdicts serve as reinforcement-learning rewards, inaccurate judgments may reinforce undesirable behavior. We introduce Sauron, an evidence-first agentic verifier that examines state and trace, independently validating proposed claims against their sources before using them to support a verdict. Bounded queries and shared investigation of related criteria reduce context use and redundant work, while evidence receipts make judgments auditable. Across 2,586 human-labeled criteria from 100 personal-assistant episodes, Sauron achieves the highest F1 score among five verifiers on all eight tested backbones. On this benchmark, it also costs less than Gandalf, the strongest baseline, in every matched comparison. With GPT-5.6 Sol, Sauron achieves 0.96 F1 compared to Gandalf's 0.92 at 23% lower cost, with larger gains on process criteria. In a paired fault-injection suite covering 13 families, Sauron detects 97% of injected failures while flagging only 2% of clean controls, outperforming baselines given the full execution trace. On TRAIL, which evaluates detection and localization of real agent errors, Sauron leads both Gandalf and the benchmark's judge on all three quality metrics at matched backbones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.