acceptodds
Under review as a conference paper at ICLR 2027

Imagined Action Sequence Guidance for Constrained Policy Optimization in Goal-Conditioned Reinforcement Learning

Abstract

Goal-conditioned reinforcement learning trains a policy to guide an agent toward a desired goal.Sparse rewards make useful and blocked actions difficult to distinguish near obstacles.Continuation Guided Constrained Policy Optimization (CGCPO) is proposed to combine imagined sequence guidance with a separate comparison of current action values.Candidate and reference actions condition a generator learned from recorded trajectories, and a sequence critic estimates the returns of the generated action blocks.These blocks are evaluated without intermediate feedback, whereas the deployed policy selects a new action after every observation.An independently trained primitive critic therefore compares the actual candidate and reference actions followed by the same reference policy.The policy objective favors a larger sequence value gap and penalizes primitive value gaps below a chosen margin.The trained policy selects one action per observation during execution.Under fixed policies and a fixed generator, the associated Bellman evaluations converge to their unique fixed points. The candidate policy's optimality gap is bounded by the sequence and primitive estimation errors, the generated block discrepancy, the candidate advantage spread and the remaining improvement in the policy score.Experiments on Fetch and PointMaze show the highest late success among the compared methods on the larger maze layouts and gains over baselines on obstacle manipulation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.