Deep Reinforcement Learning with Incomplete Symbolic Knowledge in Parametrized Action Spaces
Abstract
Specifying a complete symbolic model of a decision-making domain is difficult, tedious, and error-prone. However, even incomplete domain knowledge can constrain decision spaces without providing a complete domain model. In this paper, we study how Deep Reinforcement Learning can use such incomplete symbolic domain knowledge in parametrized action spaces, where each decision combines a discrete action with continuous parameters. We introduce our novel Knowledge- and Gradient-guided Reinforcement Learning (KGRL) algorithm that leverages on incomplete symbolic domain knowledge to reduce training time and improve safety during exploration. KGRL uses a Datalog knowledge base that encodes action restrictions and parameter bounds, which can be instantiated e.g. from domain expert knowledge. At each decision, the knowledge base is evaluated with a symbolic state abstraction, deriving state-specific action constraints and parameter bounds, while a gradient-based refinement searches the permitted parameter space for optimal parameter values. The symbolic derivations of the knowledge base additionally provide local traces that explain the imposed restrictions. Across six evaluation tasks in four different environments, KGRL achieves higher returns faster than five RL baselines for parametrized action spaces. Ablations show that including the incomplete domain knowledge accounts for substantial performance gains, whereas the benefit of the gradient guidance remains task-dependent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.