acceptodds
Under review as a conference paper at ICLR 2027

Differentiable Endogenous Neuro-Symbolic Safe Explo ration: A Log-Sum-Exp Barrier for Ppo-Lagrangian Constrained Reinforcement Learning

Abstract

Constrained policy optimisation promises provable safety, but in practice two pathologies conspire to make deployment painful. (i) The Lagrangian dual climbs the wrong ridge: with a hard-cost critic the actor treats con straint violations as sparse credit-assignment noise and oscillates between aspiration and capitulation. (ii) Conjunctions of safety predicates (e.g. “stay below the joint limit and avoid the obstacle and keep end-effector speed below the ISO 15066 limit”) collapse to whichever margin is most violated—Gödel t-norms hit a sparse gradient path, product t-norms van ish, Łukasiewicz t-norms have a non-strict dead-zone—so none of the three gives the dense, informative gradient that on-policy Ppo needs. We in troduce DENSSE (Densse), a differentiable endogenous neuro-symbolic en forcer. Margin functions are endogenous: each predicate’s effective argu ment is its physical residual plus a learned modulation ρtanh ε(s) that shares a latent code z (s) with the actor. Their conjunction uses the log sum-exp (LSE)conorm Cβ(m) = −(1/β)log ∑k exp(−βmk),which is sand wiched by mink mk and mink mk − log K/β and has fully dense, analytic Jacobian. We prove a certified-safety theorem: if Cβ(mθ) ≥ log K/β + ρ then mink ψk ≥ 0, i.e. every predicate is satisfied. An adaptive temperature controller regulates the gradient entropy H of the conorm by 1-D feedback on β, keeping the gradient useful as the residual distribution sharpens dur ing training. We instantiate DENSSE on the PPO-Lagrangian skeleton (Ppo-Lag), and evaluate on a self-contained DENSSE-Bench that includes a 2-link planar manipulator with full rigid-body dynamics (SafeReach), a double-integrator point-mass with dynamic obstacles (SafeNav), and a language-conditioned navigation task (SafeVLA-Lite) that pairs hashed n gram language features with the backbone. Across 3 seeds per (method, environment) cell at a fixed 5 ×105-step budget, the relative ordering of methods depends on the environment: DENSSE has the lowest seed vari ance on SafeReach, ties on SafeReach’s SR with the product-t-norm com piler at 0.500, sits third on SafeNav behind reward shaping (0 .750) and the Lükasiewicz compiler (0 .667), and is fourth on SafeVLA-Lite where the Gödel compiler (0 .396) and reward shaping (0 .271) lead. The ablation panel shows that disabling any single component of DENSSE either desta bilises the gradient (β too cold / uncertified budget off) or removes the shared-latent safety signal (ρ = 0). The certified-safety bound is verified empirically as a Monte-Carlo certificate in the appendix.1

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.