Shallow Networks, Distant Goals: Competence-Aware Local Planning for Online GCRL
Abstract
Learning to reach distant goals is a central challenge in online goal-conditioned reinforcement learning (GCRL). Recent advances in contrastive reinforcement learning (CRL) demonstrate the benefits of scaling network depth, alongside increased computational demands. We investigate how far shallow networks can go with a suitable planning structure. To this end, we introduce Competence-Aware Local Planning (CALP), a simple approach that constructs long-horizon behavior through repeated local decisions. CALP searches locally in goal space, using learned execution competence to identify feasible subgoals and a lower-quantile temporal-distance estimate to guide selection toward the final goal. A shared controller executes these subgoals, with all components learned jointly from online interaction. Across eight navigation and manipulation tasks in JaxGCRL, CALP improves goal-reaching performance over other online GCRL baselines using two-hidden-layer networks and a single local selection step per planning decision, with the largest gains on obstacle-interaction tasks. These results demonstrate the effectiveness of local goal-space planning for long-horizon goal reaching with shallow networks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.