Language Operators for Grounded Reward Interfaces
Abstract
Reward machines expose temporal reward structure to reinforcement learning (RL), but they are usually finite-state, grounded by hand in symbolic propositions, and learned or constructed one task at a time. These assumptions become restrictive when reward-relevant behavior requires memory beyond finite state, when symbolic task progress must be connected to continuous observations, or when an agent must solve a family of related task specifications rather than a single fixed task. We propose language operators: operators that learn maps from formal task specifications to grounded reward interfaces for RL. Given a specification, the interface predicts valid next events, task completion, and grounded targets from a trajectory prefix and the current observation. We learn these interfaces jointly across related tasks using supervision from formal monitors, and then keep the operator fixed during policy learning. The same parameters define reward-interface behavior for new specifications, allowing task descriptions to change without learning a separate reward model. We instantiate this framework with recurrent and memory-augmented models and evaluate it on BabyAI, WaterWorld, and controlled counting and nesting tasks. Across controlled counting and nesting tasks, WaterWorld, and BabyAI, we compare language operators with learned task-conditioned machines, exact symbolic references with learned grounding, and task-conditioned or recurrent RL baselines. Controlled numerical experiments further demonstrate generalization to task parameters beyond the training range, showing how shared reward-interface learning can support both new task compositions and longer executions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.