acceptodds
Under review as a conference paper at ICLR 2027

Interaction Collapses Safety: On the Query-Access Limits of Jailbreak-Robust Generation

Abstract

Recent theoretical work has formalized a trade-off between consistency and breadth in language generation from positive examples. Deployed language models are, in addition, queried: users choose prompts and observe outputs, and jailbreaking uses this access. We develop a query-access theory of jailbreak robustness, in which a generator is trained from positive data and accessed through black-box sampling. We show that a generator that is both broad and sufficiently robust yields, from samples on a fixed batch of prompts, a membership estimator for the target language. Our main result, the Query-Access Impossibility Theorem, follows: for any class of prefix-closed languages (violations cannot be repaired once emitted) that is not statistically identifiable from positive data, no generator is both uniformly robust and uniformly detectably broad once its worst-case error is small relative to its detectable breadth. Adaptivity is unnecessary for this impossibility, but it can substantially reduce the cost of discovering the boundary. All properties are quantified over the random training sample, and an explicit rate sequence and an explicit non-identifiable class show that the hypotheses hold simultaneously. Three consequences follow: robust generators lose detectable breadth, broad generators leak detectably in the worst case, and requiring average-case consistency adds no constraint. Exact or noisy membership feedback removes the barrier, and preference feedback does so exactly when its reward carries information about the boundary. The message is that a sufficiently broad generator reveals, through its own samples, the boundary it is meant to protect: when positive data do not identify that boundary, the tension between helpfulness and jailbreak robustness is information-theoretic and cannot be removed by better training on the same data, but can be removed by feedback that certifies the boundary.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.