COMBAT: Cross-Model Jailbreaking via Depth-Interpolated Activation Transfer
Abstract
Extracting intervention directions from the target model is a natural approach to representation-level jailbreaking. Target-native intervention directions already lie in the target representation space, giving them a natural compatibility advantage. We investigate whether refusal and harm-detection directions transferred from another model through representation alignment can nevertheless produce interventions that are as effective as, or even more effective than, target-native ones under the same search budget. We propose rss-odel Jailreaking via Depth-Interpolated ctivation ransfer (), a layerwise representation alignment framework that constructs banks of refusal and harm-detection directions in a source model, maps them into a target hidden space through depth-interpolated bridges that accommodate unequal model depths, identifies behaviorally effective configurations via constrained two-phase selection, and deploys them through localized interventions with a triangular layer taper. We evaluate 20 ordered source–target pairs from five aligned language models. Held-out state reconstruction assesses the learned cross-model correspondence, while equal-budget target-native controls evaluate the effectiveness of cross-model direction construction. Across four primary benchmarks, Full COMBAT achieves a micro-average GPT-judged attack success rate of 96.2%, compared with 88.7% for the equal-budget target-native control and 20.3% for unmodified targets. Further analyses show that refusal and harm-detection directions make complementary contributions, with the most effective transfers from narrow source-layer bands and trailing source direction-extraction positions immediately preceding assistant generation. The bridges fit held-out states well, and their normalized nonnegative bridge-cosine profiles decay away from selected anchors. Cross-model representation alignment thus informs the construction of effective target-side jailbreak interventions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.