acceptodds
Under review as a conference paper at ICLR 2027

Hidden Policy: The Risk Behind Risks

Abstract

Frontier models' safety and security are usually assessed through *what a model can do* and *what it has done*. Deployment decisions, however, depend on a harder question: *what will the model do?* We argue that answering this question requires identifying the *policy* that governs behavior across interaction conditions. Specifically, we define a ***hidden policy*** relative to an observer's interaction regime: a policy can be behaviorally indistinguishable from a reference policy within the regimes available to observation, yet systematically diverge under another reachable regime. This exposes a *policy-identification* challenge: behavioral evidence identifies the underlying policy only up to an ***observational equivalence class***, leaving policies that agree under the observed regime but diverge as supervision, incentives, access, authentication, or other interaction conditions change. To study this problem empirically, we construct controlled *sandbagging* policies that vary both where underperformance is expressed and how it is realized. Within construction-adjacent regimes, the outcome models closely follow the intended policy. As the observational regime expands, however, the learned policy proves less stable than expected. **Surprisingly, even such a brittle hidden policy can be difficult to eliminate.** Behavioral recovery after intervention can remain condition-specific, coexist with degradation outside the hidden regimes, or redistribute the dangerous behavior. We find that establishing the *construction* and *removal* of a hidden policy are both *policy-identification* problems across observation regimes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.