PICLAlign: Safety Alignment for Parametric In-Context Learning
Abstract
Parametric in-context learning (PICL) reduces the inference costs of in-context learning (ICL) by using a trained hypernetwork to convert document contexts into reusable LoRA adapters. Despite their promising results, across three PICL-integrated LLM systems, we find that internalizing benign documents through PICL significantly weakens safety alignment, whereas supplying the same documents through ICL causes little or no degradation. Existing alignment methods do not fix this problem because PICL introduces a distinct pattern of safety erosion (e.g., safety erosion occurs in earlier layers than the typical fine-tuned LoRAs) and additionally requires safety to be achieved across LoRAs generated from diverse documents, including unseen ones. In response, we propose PICLAlign, a unified framework that simulates PICL during alignment to better account for the safety-compromising effects of PICL-generated LoRAs. The framework supports both post-generation and generation-phase intervention: PICLAlign-L trains a safety LoRA to compose with each generated LoRA, while PICLAlign-H tunes the hypernetwork to generate LoRAs that enhance safety by construction. We further introduce Doc-DRO to improve robustness to atypical erosion patterns from unseen documents by optimizing a theoretically grounded worst-case loss over a neighborhood of document distributions. Experiments on systems integrated with two mainstream PICL frameworks, D2L and SHINE, show that PICLAlign achieves state-of-the-art safety alignment while preserving utility, reducing ASR from 35% to 0% under the D2L setting and from 56% to 0.4% under the SHINE setting. Our work reveals unexplored risks in PICL and motivates future research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.