acceptodds
Under review as a conference paper at ICLR 2027

Judging by Evidence: Typed Evidence Acquisition for Agent-as-a-Judge

Abstract

Agent-as-a-Judge, which actively interacts with environments and tools to gather verifiable evidence, has emerged as a promising alternative to rule-based verifiers and LLM-as-a-Judge for evaluating long-horizon agents and for providing rewards in agentic RL. Existing methods treat judging as a reading task, either feeding an entire trajectory to a single model or retrieving the spans most similar to each criterion. However, both break down on long trajectories, where decisive evidence is buried among tens of thousands of tokens, absence can never be established by retrieval, and state changes that the trajectory never records remain unverifiable. To address this, we propose **ATTEST**, a framework that recasts Agent-as-a-Judge from a reading task into a budgeted investigation, acquiring evidence separately for each criterion according to its logical type. Specifically, a Criterion Typer (CT) first determines what evidence each criterion requires, and a Budgeted Investigator (BI) acquires it through targeted retrieval for existential criteria, coverage scans that establish absence for universal ones, or active environment testing for stateful ones, allocating the budget by uncertainty. In turn, an Evidence Adjudicator (EA) grounds each verdict in provenance-ranked evidence rather than the agent's own narration, and abstains when evidence is insufficient. To diagnose where judges actually fail, we further construct **EviProbe**, a diagnostic benchmark that annotates the evidence supporting each criterion and pairs every instance with minimally edited counterfactuals, separating evidence-retrieval failures from verdict failures. Across three benchmarks, ATTEST outperforms both LLM-as-a-Judge and prior agentic judges under matched interaction budgets, suggesting that how evidence is acquired, rather than how much a judge reasons, is the primary lever for reliable agentic evaluation. Code is available at: https://anonymous.4open.science/status/Judging-by-Evidence-EF27.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.