HILT: Hazard-Informed Localized Teaching for Safer Language Models
Abstract
Safety training for tool-using language models must correct hazardous decisions while retaining useful task behavior. Trajectory-level feedback identifies unsafe episodes but leaves unclear where corrective learning should begin and what evidence should guide it. We introduce HILT (Hazard-Informed Localized Teaching), a framework that couples the student's starting context with verifier evidence supplied to a privileged teacher. This connects the decision being corrected with evidence of its harmful consequences, enabling focused supervision of safer continuations. On-policy self-distillation transfers the teacher's guidance through student-sampled text continuations into a policy that requires no safety verifier at inference. Across five models spanning three families, HILT achieves the lowest observed average attack success rate among the compared methods in four of five settings. Compared with whole-trajectory self-distillation, it achieves lower observed average attack success and over-refusal in all five settings, with absolute over-refusal reductions of 6.09%–23.60%. Aggregate reasoning accuracy remains within 1% of each base model in absolute terms.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.