KeyedRoPE: Structural Capability Control through Key-Conditioned Attention Geometry
Abstract
Can access to a language model's learned capabilities be made contingent on its internal attention geometry rather than an external textual credential? We investigate this question through KeyedRoPE, a structural capability-control mechanism that conditions rotary attention geometry on a compact nontextual key. KeyedRoPE transforms token positions and reconstructs RoPE phases, making useful model execution dependent on an authorized attention geometry. Across four language models from 410M to 8B parameters and four downstream tasks, protected models largely retain authorized utility while performance substantially degrades under standard RoPE and held-out incorrect keys. Under adaptive white-box evaluation, gradient-based geometry search yields partial capability recovery, while fine-tuning recovery increases with task-specific supervision and is slower than Password-Locked Models at limited and intermediate data budgets. In our Qwen3-4B Mobile Action evaluation, full-data adaptation recovers 96.2% of the pre-attack authorized utility and 92.13% of the original task-adapted utility under standard RoPE. These results do not establish that full-data retraining is necessary for recovery. Instead, they show that KeyedRoPE transforms direct checkpoint reuse into an explicit recovery and re-adaptation problem whose difficulty can be measured under specified attacker budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.