acceptodds
Under review as a conference paper at ICLR 2027

Learning to Follow Instructions from Imperfect Guidance: Reference-Guided Reinforcement Learning with Demonstration Anchoring

Abstract

Planners, experts, and language models can guide instruction-following reinforcement learning (RL), but may receive inaccurate state descriptions or be impractical to query at deployment. To address this problem, we initialize policies by imitation and fine-tune them with rewards, temporary action-probability targets from a frozen reference on learner-visited states and continuing demonstration supervision (anchoring). We train policies that learn from imperfect guidance and act without it. In BabyAI, with left/right-swapped inputs to an online language guide, anchoring raises mean final success from 18.3% to 27.7% (five seeds, one initializer). Separately, across five independent BabyAI datasets and model fits, the framework improves time-averaged success by 6.5 percentage points over anchoring alone and 8.1 over the unchanged initial policy. Exchanging the same targets among states sharing an instruction lowers time-averaged success by 4.6 points relative to own-state guidance. Another BabyAI study restricts exchanges to targets sharing the most probable action, reducing this gap from 5.1 to 0.7 points; further benefit from exact own-state probabilities remains unresolved. In CrafText, own-state guidance beats reassignment, uniform targets and anchoring alone, but not the initial policy on main-study time-averaged success, despite higher final success.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.