Trust Is a Process, Not Just a Score: Agentic Verification for Test-Time Scaling
Abstract
While test-time scaling improves large language models (LLMs) by generating multiple candidates, verification is often treated as a terminal scoring problem, separate from the acquisition and maintenance of evidence. In this paper, we explore verification as an interleaved process of gathering observations and updating a validity-aware state: execution probes, knowledge-base traversals, and model adjudication distinguish candidates, while the state tracks dependence, supersession, disclosure, and repair. We instantiate this approach, named GEVA (Governed Evidence Verification and Audit), as a deterministic state–action–observation loop over a typed claim–evidence graph, with certificate-based audits governing evidence producers. Across seven regimes spanning code, text-to-SQL, knowledge-graph question answering, and multiple-choice reasoning, GEVA improves over greedy decoding throughout. On MetaQA-3hop, it reaches , points above common-pool Fixed V; matched-evidence controls show that this recovery comes from executed retrieval rather than graph storage alone. On a disjoint held-out split, producer audit restores accuracy points by withdrawing a harmful judge. Event replay further reproduces full recomputation with nearly fewer claim assessments. These results demonstrate the benefits of jointly acquiring, maintaining, and governing verification evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.