acceptodds
Under review as a conference paper at ICLR 2027

CPO-RL: COT-POLICY CO-OPTIMIZATION FOR EMBODIED REINFORCEMENT LEARNING

Abstract

Reinforcement learning (RL) can adapt Vision-Language-Action (VLA) models through interaction, but reward feedback alone does not explicitly identify the contact regions needed for precise manipulation. We propose CPO-RL, a CoT–Policy Co-Optimization framework that jointly trains a contact-grounded intermediate representation and a continuous control policy. CPO-RL uses dynamic 2D contact masks as supervision for an embodied Chain-of-Thought (CoT) latent. An input-dependent contact token conditions action prediction, while an auxiliary decoder combines this latent with RGB features to predict the contact mask. Contact-SFT first initializes the policy and the contact representation from demonstrations. During online learning, PPO uses all collected trajectories, while valid contact targets derived from successful simulated interactions supervise the auxiliary branch and its shared policy representation. This couples policy optimization with refinement of contact localization without requiring predicted masks or the auxiliary decoder at deployment. On four ManiSkill manipulation tasks, CPO-RL achieves 94.2% average success, improving over the official RLinf baseline by 19.3 percentage points and reducing mean episode length from 68.5 to 52.7 steps. A controlled component ablation shows that contact supervision contributes beyond adding the latent token alone. After sim-real co-training, deployment on a physical Franka robot yields 66.9% average success across original and distractor settings, compared with 29.4% for SFT. These results support contact-grounded representation learning as a practical complement to reward-based VLA post-training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.