What Can AI Hide? A Framework and Toolkit for Evaluating Hidden Capabilities
Abstract
An AI system could hide an ability while being tested and continue to use it when answering ordinary requests. Yet those responses may contain information it withholds during testing. We introduce a framework and toolkit for examining what an evaluator can recover from those responses. To measure how well the system handles ordinary requests, we use a score defined for the task, such as the proportion of answers it gets right. We then ask whether the system can keep the tested answers hidden while meeting a required score on those ordinary requests. Our framework describes when both are possible and when staying hidden requires accepting a lower score. To apply the framework, an evaluator specifies what information it can inspect and which ways of hiding the evaluation covers. We provide tools for choosing requests, scoring responses and determining what conclusions the results support. For repeated testing, we give an upper limit on the chance of rejecting a hiding explanation by mistake, provided the assumptions about sampling and responses hold. Our experiments show why recovering answers and preserving ordinary performance must be checked together. In one model study, an evaluation appeared successful mainly because it recovered answers that the instruction to hide had barely changed. Together, the framework and toolkit provide a way to evaluate what an AI system can keep hidden while still answering ordinary requests to a required standard.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.