From Causal Graphs to Markovian States: Multi-Order State Exposure for Deep RL
Abstract
Online reinforcement learning (RL) relies on the Markov property for guaranteed performance, but real-world applications often lack well-defined states given raw observed variables. While causal RL has attracted growing interest, existing work typically assumes Markovian states are provided and uses causality only to accelerate learning. This leaves a fundamental gap: given a longitudinal causal graph over observed variables, how does one construct MDP states that provably satisfy the Markov property? We provide a procedure that constructs a provably minimal such state.In deep RL, the minimal representation alone empirically fails to improve performance, indicating that neural networks cannot automatically exploit Markovian minimality. We therefore propose MOSE (Multi-Order State Exposure), which exposes a single shared Q-function to history-based state constructions of orders 1 through simultaneously. An attention-based selection layer scores the dimensions of this combined input by their estimated relevance to value-function estimation and selects the high-scoring ones. Causal MOSE further constrains the selection layer to retain every variable of the minimal state, guaranteeing sufficiency while leaving the model free to admit useful redundancy. MOSE consistently outperforms both the minimal state construction and single-window policies on common benchmarks and synthetic datasets. Causal MOSE can further improve performance. Our results establish a core principle for causal deep RL: minimal sufficiency is not enough. Controlled redundancy gives the model flexibility to learn its state representation end-to-end. Causal state information provides structural guidance that can make this learning more efficient.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.