acceptodds
Under review as a conference paper at ICLR 2027

Beyond Refusal: Policy-Grounded On-Policy Distillation for Safety Decisions and Recovery

Abstract

Safety distillation can teach a language model to refuse without teaching it the policy distinction behind the refusal—or how to recover once generation has entered a harmful trajectory. We study this gap in on-policy distillation (OPD) through representation readouts, trajectory diagnostics, and budget-matched interventions. Rule and boundary generalization lag behind refusal, while recovery states become less frequent in student rollouts as refusals emerge; explicit policy supervision improves boundary decisions, and targeted state coverage further improves recovery. Motivated by these findings, we introduce DEEPEN, a policy-grounded OPD framework that makes safety judgments explicit and restores access to decision-critical states. DEEPEN predicts structured safety labels to condition generation, trains on boundary pairs and harmful-prefix starts, and uses brief teacher handoffs triggered by policy errors to supply a corrective segment before returning control to the student. Overlap-aware token-level distillation reinforces local corrections, while handoffs anneal to zero for student-only deployment. Across four preselected representatives from distinct model families under matched response-token budgets, DEEPEN improves mean boundary score (BDA) by 6.6 percentage points and useful recovery by 18.1 points over an adapted PolicyAlign baseline; the largest task-wise capability decrease is 4.5 points relative to the initial student. These results support treating safety distillation as a joint problem of policy-decision learning and safety-critical state coverage, rather than refusal imitation alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.