Learning Reachability Rewards for Temporal-Logic Reinforcement Learning
Abstract
Reinforcement learning with Linear Temporal Logic (LTL) specifications faces severe credit-assignment challenges when learning infinite-horizon Büchi objectives. We propose RACE (Reachability-Aware Credit Estimation), a framework that uses accepting-set reachability as a surrogate for learning recurrent acceptance. RACE trains a discriminator on trajectories collected by the agent to predict whether the rollout continuation from a sampled state reaches acceptance. Its predictions provide dense, state-dependent rewards for policy learning. We extend this construction to intermediate stages of the limit-deterministic Büchi automaton (LDBA), allowing partially successful trajectories to provide useful feedback even when they do not reach acceptance. Automaton-induced potentials complement these learned rewards by guiding exploration. Our analysis establishes a scaled lower bound on an infinite-time reachability objective and provides conditional connections to Büchi satisfaction. Empirically, RACE improves exploration efficiency and task performance across continuous-control domains with complex LTL objectives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.