InEx: Intrinsic and Extrinsic Evidence for Trustworthy Evaluation of Latently Verifiable Tasks
Abstract
Many real-world agent tasks are latently verifiable: a correct answer or success condition exists in principle, but the corresponding ground truth is not cheaply available during evaluation. Reliable evaluation must therefore determine whether an agent's claimed outcome is supported by evidence rather than merely judge the plausibility of its final answer. We introduce InEx, an evidence-based evaluator that combines two complementary sources: intrinsic evidence already contained in the submitted trajectory and extrinsic evidence acquired through additional environment interaction. InEx-I reconstructs task-relevant intrinsic evidence by inspecting local observation-action transitions and determines whether the resulting evidence chain supports the claimed outcome. When intrinsic evidence alone is insufficient, InEx-E selectively interacts with the environment to acquire the missing extrinsic evidence. A scaffolded handoff carries the structured intrinsic analysis into extrinsic verification so that the two components form a continuous evidence chain. We implement InEx-I by specializing an open vision-language model through supervised fine-tuning, reinforcement learning, and confidence-aware learning from imperfect supervision, and implement InEx-E with a frontier model capable of environment interaction. We evaluate InEx on a diverse suite of agent benchmarks spanning web, mobile, and desktop interaction. InEx consistently achieves the highest evaluator accuracy among the compared methods and produces more reliable rewards that improve downstream agent training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.