When Does Cross-Examination Beat Self-Assessment? A Text-Surface Visibility Law for Training-Free Agent Trustworthiness
Abstract
Detecting when an LLM agent is about to fail—without ground truth or extra training—is central to deploying agents safely. Two training-free signals compete: self-assessment (the acting model reports its own confidence) and cross-examination (a separate model audits the trajectory). We ask when one beats the other, and find that neither wins universally. Across five agentic settings spanning conversational tool use (τ-bench retail/airline), step-level failure attribution (Who&When), and code generation (BigCodeBench), we establish a monotone visibility law: the net benefit of cross-examination over self-assessment is determined not by whether the task has an objective notion of correctness, but by whether the error signal is exposed at the text surface readable by the auditor. Ordering domains by a semi-quantitative visibility rubric yields a strictly monotone gain curve, from a significant negative gain in low-visibility conversational trajectories (−0.156, P=.02) through a near-zero point in code—a high-objectivity but low-visibility counter-example (+0.024, ns)—up to a significant positive gain in action-rich retail dialogues (+0.305, P=1.0). We further show that naive multi-judge aggregation does not beat the strongest single judge, and that inter-judge agreement itself tracks visibility. Finally, we give a first partial solution: a deploy-time observable proxy (trajectory depth) routes each domain to the better signal, reaching 97% of an oracle's within-domain AUROC and significantly outperforming always-self (P=.999). All findings are training-free and pure-API.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.