acceptodds
Under review as a conference paper at ICLR 2027

Why Does Long-Context LLM Safety Break? A Positional Encoding Perspective

Abstract

Large Language Models (LLMs) have recently demonstrated strong long-context capabilities, enabling them to process and reason over increasingly extensive inputs. However, this capability also exposes them to many-shot jailbreaking (MSJ), in which long sequences of adversarial demonstrations can weaken refusal behavior on harmful requests. We find that the RoPE base is closely associated with this vulnerability: increasing it substantially lowers attack success, but directly doing so also reduces long-context utility. An analysis of RoPE interactions shows how input length and the base jointly shape rotary phases across token distances. This suggests a way to use positional variation during training without changing the model's inference setting. We introduce MFA, a multi-frequency alignment framework that samples RoPE bases during GRPO training and strengthens safety updates for longer prompts. At inference, the model retains its native RoPE base. On Llama-3.1-8B-Instruct under 256-shot MSJ, MFA reduces attack success rate from 65.5% to 8.8%, compared with 28.0% for GRPO trained at a single base. Across the evaluated models and attack lengths, MFA improves many-shot safety while maintaining performance on general and long-context tasks. Code is available at https://anonymous.4open.science/r/mfa-EBA4

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.