acceptodds
Under review as a conference paper at ICLR 2027

From Shallow to Deep: Pinning Semantic Intent via Causal GRPO

Abstract

Safety-aligned large language models (LLMs) remain surprisingly vulnerable to adversarial prefixes and prefilling attacks, where a short compliant context can steer an otherwise refusing model toward harmful continuation. We investigate this failure from a representation-level perspective and observe a phenomenon we term Semantic Representation Decay: harm-relevant information that is readily accessible at the end of a malicious query becomes substantially less stable as generation proceeds under adversarially induced contexts. Motivated by this observation, we propose Two-Stage Causal-GRPO (TSC-GRPO), a post-training framework for robust safety recovery. In Stage 1, we construct controlled contextual interventions that preserve the underlying harmful semantics while varying prefixes, adversarial perturbations, and partial generations. A lightweight probe is trained to learn a context-invariant harm representation, while remaining sensitive to genuine transitions from harmful continuation to refusal. In Stage 2, the frozen representation model serves as a semantic scorer during GRPO. By measuring how long each sampled continuation remains aligned with a fixed harmful-semantic anchor, we define a dense cumulative trajectory penalty that assigns higher reward to earlier safety pivots. Experiments show that TSC-GRPO significantly outperforms baselines in defending against jailbreak attacks while preserving general utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.