Beyond a Single Safety Horizon: Multi-Horizon Policy Coordination for Offline Safe Reinforcement Learning
Abstract
Safe offline reinforcement learning aims to maximize task performance under safety constraints using only static datasets. Existing approaches typically deploy a single policy, collapsing potentially different temporal safety preferences into one behavior. However, safety violations can unfold over substantially different time scales, such that changing the safety horizon alters which future hazards are visible when evaluating an action and can induce distinct safety-oriented behaviors. This motivates us to exploit temporal safety specialization rather than compressing multiple safety scales into a single policy. We propose **Multi-Horizon Safety Policy Coordination (MSPC)**, which learns and coordinates a family of safety-first policies specialized to different temporal horizons. At the low level, safety supervision over multiple future horizons is used to extract complementary, behavior-supported policies from the same offline dataset. At the high level, shared reward and cost critics evaluate their candidate actions, and a residual-budget coordinator selects among them based on the remaining safety allowance. We further use policy disagreement to avoid redundant coordination when candidate actions remain similar. Experiments on 18 tasks from two offline safe RL benchmark suites show that the evaluated fixed specialists exhibit non-nested safety outcomes, highlighting the need for adaptive coordination among complementary temporal safety behaviors. MSPC further achieves favorable reward–safety performance against representative offline safe RL baselines, demonstrating the effectiveness of temporal safety specialization and budget-aware action-level coordination.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.