Following by Contrasting: Contrastive Self-Distillation for System-Prompt Adherence
Abstract
System prompts specify the roles, constraints, and preferences that large language models (LLMs) should follow. However, models often fail to reliably respect these requirements, particularly when they conflict with user requests or learned behavioral tendencies. Existing approaches to improving system-prompt adherence commonly rely on external behavioral supervision or additional computation during decoding. To address this, we propose Contrastive Self-Distillation for System-prompt adherence (CoSSy), a self-distillation framework that learns from the model's own prediction shift under system conditioning. Specifically, CoSSy contrasts predictions with and without the system prompt, amplifies the resulting shift, and distills it into a student conditioned on the system prompt, thereby learning to more effectively use the provided system prompt. This eliminates the need for external responses or preference labels while retaining standard autoregressive decoding at inference time. Across three LLMs on general system-message following and challenging system-prompt adherence benchmarks, CoSSy consistently outperforms competitive baselines. For example, CoSSy improves over the base models by 8.6% relative on average on Multifaceted-Bench and by 49.7% on IFBench Strict. Our analyses further reveal that CoSSy internalizes the supervision as an adaptive, state-dependent contrast rather than a fixed decoding-time intervention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.