acceptodds
Under review as a conference paper at ICLR 2027

When Confidence Is Not Enough: Action-Contrastive Self-Models under Capability Drift

Abstract

A calibrated agent may estimate its probability of success accurately and still choose the wrong recovery. Retrieval, verification, replanning, escalation, and abstention repair different failures, so scalar confidence removes the direction needed for action choice. We formulate self-knowledge as an action-contrast surface: the change in terminal value induced by each available recovery branch. SCOPE factors this surface through learned operational response modes, aligns their action–diagnostic signatures after a model, tool, access, or context change, and routes by predicted net value. Under conditional response homogeneity, is sufficient for one-step recovery choice. In a multi-regime linearization, sharing modes across regimes changes the stylized parameter count from to , while signature separation gives a condition for transporting persistent modes across versions. We evaluate these claims with real local-model and tool executions: every legal branch is run from the same reset pre-action state, terminal outcomes are scored by task-native evaluators, and only one logged branch outcome is exposed to the learner. Across 6 open-weight checkpoints and six transition families, action conditioning raises mean normalized utility from to , generic rank-matched sharing reaches , and SCOPE reaches . SCOPE further reduces held-out intervention-gain error by – relative to the rank-matched control. Direct-AV catches up with abundant feedback and achieves higher utility when new response mechanisms or noisy diagnostics break the shared geometry.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.