acceptodds
Under review as a conference paper at ICLR 2027

GEAR: Generative Efficient Gap Amplification for Goal-Conditioned Reinforcement Learning

Abstract

In goal-conditioned reinforcement learning (GCRL), trajectories that successfully reach the same goal can differ substantially in behavioral efficiency. Existing work on goal relabeling and generative sampling improves data reuse, but pays limited attention to how efficiently goals are reached, potentially reinforcing redundant behaviors when learning from successful experience. We propose GEAR (Generative Efficient Gap Amplification for GCRL), a generative framework that guides policy learning by contrasting efficient and redundant goal distributions. GEAR trains two diffusion-based goal generators through weighted sampling to reflect trajectory efficiency while retaining useful information from longer successful trajectories. The efficient generator emphasizes goals from trajectories that successfully reach the goal in fewer steps; the redundant generator emphasizes goals from longer successful and failed trajectories and is adversarially trained to produce plausible distractors. A dual-critic architecture evaluates the same state-action pair under both goal conditions, and the resulting value gap serves as an auxiliary contrastive signal that encourages the actor to favor actions that better support efficient goals while providing less support for redundant goals. Experiments on DMC, ManiSkill2, and Habitat show improved task performance and sample efficiency. On a physical robotic arm, GEAR achieves 90%, 75%, and 80% success on object lifting, block stacking, and target grasping among distractors, respectively, outperforming all evaluated baselines on each task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.