acceptodds
Under review as a conference paper at ICLR 2027

Compositional Language-Instructed Reinforcement Learning

Abstract

A compositional natural-language instruction decomposes a sparse-reward task into an ordered sequence of subgoal instructions, without providing a criterion for when to advance. Most prior approaches learn a completion signal and deliver it to the agent as auxiliary reward. We present GLEAN (Grounded Language-Effect ANchoring), which learns a completion detector from segmented offline demonstrations and uses it to advance the policy through the subgoals. The detector summarizes the egocentric observations so far into a filter state, which is also an input to the policy. It scores the change in that state since the previous completed subgoal against the active subgoal instruction and compares the score with a threshold calibrated on held-out demonstrations. The policy is initialized by behavior cloning and, throughout online reinforcement learning, continues imitating the demonstrations' subgoal segments while each detection advances it to the next subgoal instruction; we call this grounded subgoal anchoring. On six BabyAI levels with a matched interaction budget, GLEAN achieves higher final success than every baseline that receives a completion signal as reward on all five compositional levels, including baselines that start from a policy behavior-cloned on the same demonstrations. Providing the environment's exact completion predicates as reward still yields lower final success than GLEAN with its learned detector, so the advantage of advancement with continued imitation over reward does not come from detector errors

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.