acceptodds
Under review as a conference paper at ICLR 2027

Do Frontier Models Know What They Don't Know? A Unified Protocol for Trustworthiness Evaluation across Vision, Retrieval, and Agentic Tasks

Abstract

Frontier models are increasingly capable, yet deployment hinges on a different question: do they know when they are wrong? We argue that trustworthiness — abstaining under insufficient or corrupted evidence, flagging conflicts, and diagnosing failures — should be measured directly rather than inferred from accuracy. We introduce a unified, defect-conditioned protocol that injects controlled defects into otherwise-solvable instances and rewards models for the appropriate response (answer, abstain, or conflict-aware), with programmatic gold labels, multi-judge voting (Fleiss ), and stratified failure audits. We instantiate the protocol in three scenarios spanning modality, evidence, and action: visual claim verification (S1, shared image–claim probes), conflicting-retrieval calibration (S2, instances), and agentic trace diagnosis (S3, multi-step tasks) — all under an identical roster of six frontier models. Across all three we find that (i) self-reported confidence is a weak error detector, inversely calibrated on S2 (error-ranking AUROC on every contested slice) and only mildly informative on S1 (AUROC ), while cross-model disagreement is the stronger signal in all three scenarios and strongest under agentic action (S3 AUROC vs. self-confidence , where half the agents self-report near-maximal confidence at chance-level AUROC even when they fail); (ii) a consensus blind spot exists where models fail together on plausible-but-wrong inputs (stale and adversarial evidence), so cross-model disagreement — otherwise the strongest signal — collapses below chance; and (iii) neither introspection nor self-consistency resampling rescues this blind spot, though a lightweight evidence-critique prompt closes the adversarial half of it (misled-answer rate ) while the stale half stays open. The unified protocol exposes a shared failure axis invisible to any single-scenario benchmark: models are reliable when defects are overt and confidently wrong when defects are plausible. We release the protocol, data, and judge harness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.