Selecting Decision-Relevant Concepts in Reinforcement Learning
Abstract
Training interpretable concept-based policies requires practitioners to select which human-understandable concepts an agent should use for sequential decision-making. This process requires domain expertise, scales poorly with the number of candidate concepts, and provides no performance guarantees. To overcome this limitation, we propose the first algorithms for principled automatic concept selection in sequential decision-making. Our key insight is that concept selection can be viewed through the lens of state abstraction: a concept is decision-relevant if removing it causes the agent to confuse states that require different actions. That is, states with the same selected concept representation should share the same optimal action, preserving the decision structure of the original state space. This perspective leads to Decision-Relevant Selection (DRS), which selects a subset of concepts from a candidate set and provides performance bounds relating abstraction error to policy performance. Empirically, DRS performs comparably to manually curated concept sets and improves policy performance under test-time concept intervention across reinforcement learning benchmarks and a healthcare simulator.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.