What Can We Infer from Agent Traces? Observation–Claim Identifiability for Action-Routing Evaluation
Abstract
Agent evaluation requires checking not only recorded outputs, but also whether their interpretation preserves the intended question and whether the evidence determines its answer. We introduce an executable source-to-claim contract that distinguishes a changed question, insufficient evidence, and a conclusion requiring revalidation after its supporting inputs change. The workflow checks proposed representation corrections against the original question, acquires evidence from a declared catalogue, and revalidates the resulting answer. We evaluate it on source-linked queries using retained navigation, text-interaction, and hosted tool-use traces, with recorder-conformance checks on 10,800 execution records. A controlled asynchronous execution case shows that identical client-timeout projections can conceal different later transaction outcomes; acquiring transaction evidence resolves the ambiguity. On the tested contracts, direct SMT and finite enumeration reach the same conditional conclusions as the integrated workflow. In a controlled three-layout study of 36 retained records, selective access reads fewer, equal, or more bytes than direct access to the required evidence; a minimal directly available field can be cheaper than a selectively acquired ledger. The contribution is a verifiable path from an evaluation question to evidence and a qualified answer. Guarantees remain conditional on the supplied semantics, admissible execution space, and evidence catalogue.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.