TRAP: Auditing Target Risk After Prompt-Side Acceptance
Abstract
Prompt-side acceptance is decided before a response exists, but downstream harm also depends on the model that answers. We test whether the downstream risk meaning of a frozen prompt-side release decision survives response-target substitution. A prospective recovery policy releases the same 109 of 150 source-unsafe OR-Bench prompts to four targets, including a base-model stress target. Harmful-release mass ranges from 0.67% to 64.00%; the Mistral versus Qwen2.5 gap is 26.0 percentage points, corroborated by independent grading and a blinded human audit. A distinct direct-screen policy replicates target dependence on 1,000 sealed unsafe prompts, with a prespecified 24.3-point Ministral versus gpt-oss gap (97.5% CI [20.4,28.2]). Modern targets under the recovery policy vary less. To test the boundary of this result, we identify an input-guard policy that satisfies a prespecified Qwen3 risk-coverage contract using guards disjoint from the response evaluators. On a separate sealed population of 2,000 unsafe and 150 legitimate prompts, harmful-release mass is at most 1.10% for Qwen3 and Gemma-3 under each evaluator. Under Gemma-3 substitution, neither the predeclared 5% risk-violation criterion nor the three-point paired-margin criterion is met. Target dependence therefore need not violate a qualified contract; downstream assurance requires validation under the declared deployment contract.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.