acceptodds
Under review as a conference paper at ICLR 2027

Sabotage capability is predictable from five parameters

Abstract

Measuring AI's ability to cause harm despite control measures informs risk assessments. In existing sabotage capability evaluations, when we evaluate an attacker against a monitor in one setting, we do not know how the result generalizes to other models or settings, or what scores warrant stronger control measures. To address this limitation, we posit that model sabotage capability has latent structure that enables prediction of unmeasured outcomes, and find evidence for this hypothesis. With just three parameters per attacker and two parameters per monitor, we can predict how well an attacker AI performs against a monitor without having measured the pair. We can also translate these parameters between control settings to predict sabotage success probability in different evaluations. The predictive power of our simple model contributes towards more integrated, sample-efficient, and decision-relevant sabotage capability evaluations. We furthermore use this model to find that sabotage capability trends increase with general capabilities, as measured by ECI.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.