acceptodds
Under review as a conference paper at ICLR 2027

Trust Is a Process, Not Just a Score: Agentic Verification for Test-Time Scaling

Abstract

While test-time scaling improves large language models (LLMs) by generating multiple candidates, verification is often treated as a terminal scoring problem, separate from the acquisition and maintenance of evidence. In this paper, we explore verification as an interleaved process of gathering observations and updating a validity-aware state: execution probes, knowledge-base traversals, and model adjudication distinguish candidates, while the state tracks dependence, supersession, disclosure, and repair. We instantiate this approach, named GEVA (Governed Evidence Verification and Audit), as a deterministic state–action–observation loop over a typed claim–evidence graph, with certificate-based audits governing evidence producers. Across seven regimes spanning code, text-to-SQL, knowledge-graph question answering, and multiple-choice reasoning, GEVA improves over greedy decoding throughout. On MetaQA-3hop, it reaches , points above common-pool Fixed V; matched-evidence controls show that this recovery comes from executed retrieval rather than graph storage alone. On a disjoint held-out split, producer audit restores accuracy points by withdrawing a harmful judge. Event replay further reproduces full recomputation with nearly fewer claim assessments. These results demonstrate the benefits of jointly acquiring, maintaining, and governing verification evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.