acceptodds
Under review as a conference paper at ICLR 2027

When Does Cross-Examination Beat Self-Assessment? A Text-Surface Visibility Law for Training-Free Agent Trustworthiness

Abstract

Detecting when an LLM agent is about to fail—without ground truth or extra training—is central to deploying agents safely. Two training-free signals compete: self-assessment (the acting model reports its own confidence) and cross-examination (a separate model audits the trajectory). We ask when one beats the other, and find that neither wins universally. Across five agentic settings spanning conversational tool use (τ-bench retail/airline), step-level failure attribution (Who&When), and code generation (BigCodeBench), we establish a monotone visibility law: the net benefit of cross-examination over self-assessment is determined not by whether the task has an objective notion of correctness, but by whether the error signal is exposed at the text surface readable by the auditor. Ordering domains by a semi-quantitative visibility rubric yields a strictly monotone gain curve, from a significant negative gain in low-visibility conversational trajectories (−0.156, P=.02) through a near-zero point in code—a high-objectivity but low-visibility counter-example (+0.024, ns)—up to a significant positive gain in action-rich retail dialogues (+0.305, P=1.0). We further show that naive multi-judge aggregation does not beat the strongest single judge, and that inter-judge agreement itself tracks visibility. Finally, we give a first partial solution: a deploy-time observable proxy (trajectory depth) routes each domain to the better signal, reaching 97% of an oracle's within-domain AUROC and significantly outperforming always-self (P=.999). All findings are training-free and pure-API.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.