Rethinking Safety Reward Models in Long-Context Settings: Systematic Evaluation and Reliable Contextual Modeling
Abstract
Safety Reward Models (SRMs) evaluate whether LLM responses follow safety guidelines (e.g., harmlessness and legal compliance), which are essential for safety alignment during post-training and as guardrails for filtering non-compliant generations at test time. Current SRMs can accurately identify the safer response in single-turn QA tasks by evaluating isolated request–response pairs. As LLM usage expands into long-context agentic workflows, the safety preference associated with the same response may vary across contexts. Crucial contextual factors include whether the requester holds proper authorization, whether a sub-agent exceeds its assigned operational boundary, and whether environmental conditions (e.g., open network ports) expose systems to attack. However, the reliability of safety reward models in long-context settings remains under-studied. We therefore introduce ContextFlip, a benchmark designed to evaluate the safety preferences of SRM under long-context settings along three core dimensions: contextual preference reversal, candidate order invariance, and situational evidence retention. Across a diverse range of existing SRMs, our evaluation reveals a substantial performance gap between conventional request–response safety preference modeling and contextual reward modeling. Specifically, existing RMs are dominated by request-level priors, exhibit order bias, and progressively lose contextual control in long histories. To address this gap, we propose ContextRM, a context-aware safety reward model designed for reliable contextual safety preference. ContextRM incorporates a symmetric preference training objective to ensure candidate-order invariance and recursively compresses safety-relevant context into a memory state to preserve long-horizon situational evidence. Extensive experiments demonstrate that ContextRM achieves superior preference accuracy under long and complex situational inputs, ensuring reliable context-aware safety evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.