SafeForge: Dynamic Neuro-Symbolic Safety Rule Synthesis for Embodied Action Interception
Abstract
Embodied AI agents operating in physical environments must evaluate the safety of proposed actions before execution — a challenge that demands both formal logical guarantees and the flexibility to handle diverse object combinations unseen during design. Existing approaches rely either on rigid rule-based systems that fail on novel entity configurations, or on neural models that lack interpretability and formal safety certificates. We present , a dynamic interpretable neuro-symbolic safety guardrail framework for embodied action safety verification. SafeForge operates through three stages. First, a real-time scene graph builder dynamically models the embodied interaction environment, grounding 3D object positions and semantic attributes into soft spatial predicates via differentiable sigmoid distance kernels. Next, a differentiable neuro-symbolic reasoner performs hierarchical forward chaining with learnable rule weights, instantiating high-level safety rules into scene-grounded, concrete safety conditions tailored to the observed entity configuration. Finally, an action safety judge synthesizes the proposed action type with the entities involved to produce a binary safety decision, accompanied by a full proof trace that provides human-interpretable explanations. We evaluate SafeForge on a curated benchmark of tasks spanning procedurally generated scenes and hazard categories, with an average of action steps per sequence. Extensive experimental results demonstrate that SafeForge achieves strong detection performance while maintaining differentiability for downstream policy learning, outperforming purely neural and purely symbolic baselines in safety detection performance. The source code is available at https://anonymous.4open.science/r/SafeForge.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.