acceptodds
Under review as a conference paper at ICLR 2027

EpiPlan: Prospective Epigraph-Space Planning for Safe Multi-Agent Reinforcement Learning

Abstract

Epigraph-conditioned safe reinforcement learning trains a controller across task-budget levels, but deployment typically collapses this structure to the policy at a learned critic root. In distributed epigraph control, local critic roots can be aggregated into a public team boundary, but executing only that boundary member does not compare the joint feedback trajectories induced by nearby boundary-relative choices. We introduce Prospective Epigraph-Space Planning (EpiPlan), which treats the frozen controller as a one-dimensional family of complete team policies and selects a member at deployment without retraining or optimizing joint-action sequences. From the same augmented state, EpiPlan rolls out one-sided displacements, recomputes the public root along each branch, uses frozen task and safety critics as terminal estimates, and executes the lowest-cost predicted-feasible member while retaining its displacement. We formalize this actual-receding family and its deployment epigraph; under stated approximation and margin conditions, our analysis bounds finite-realization error, selected-policy feasibility residual, and task loss in terms of continuation, calibration, and coordinate resolution. Across 18 native-eligible frozen checkpoints in six settings drawn from five cooperative tasks, EpiPlan lowers episode unsafety relative to paired public-root execution, while every environment-mean reward loss remains within a prespecified 15% tolerance. An exact-marginal permutation control that preserves each checkpoint's offset histogram but breaks state–member pairing raises expected episode unsafety by 10.38 percentage points (p=.0001), showing that the gain depends on contextual assignment rather than offset frequencies alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.