Safety Hacking in Constrained Best-of- Inference-time Scaling
Abstract
Inference-time scaling pipelines often sample multiple outputs, filter them with a learned safety model, and return the proxy-feasible output with the highest learned reward. We show that this pipeline can fail in two stages: an imperfect safety proxy first contaminates the feasible set with unsafe outputs, and reward maximization can then amplify this residual contamination. We call this failure safety hacking, in which the selected output passes the learned filter but violates the true safety constraint. For constrained Best-of- sampling, we derive finite- bounds determined by the joint upper reward tails of safe and unsafe outputs within the proxy-feasible set. If unsafe-but-feasible outputs have the heavier tail, the probability of safety hacking approaches as grows, even when false-positive mass and average safety and reward proxy errors are arbitrarily small. We also show that policies within a bounded divergence from the proxy-feasible reference distribution admit an -independent safety-hacking bound. Coverage control limits amplification but cannot repair a contaminated feasible set: admitted unsafe outputs may still be favored, and regularized selection is not necessarily safer than constrained Best-of-. Toy and language-model experiments characterize both contamination and its reward-tail amplification, which exposes an inherent difficulty in inference-time scaling with imperfect reward and safety models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.