acceptodds
Under review as a conference paper at ICLR 2027

CAPO: Real-World Reinforcement Learning for VLA via Corrective Adjoint Policy Optimization

Abstract

Human-in-the-Loop Reinforcement Learning (HiL-RL) offers a practical route for adapting Vision-Language-Action (VLA) models through real-world interaction guided by human interventions. Yet effectively leveraging these interventions remains challenging, their corrective signals must first support informative value learning and then be translated into effective updates of the VLA's multi-step flow policy. To address these challenges, we propose **C**orrective **A**djoint **P**olicy **O**ptimization (**CAPO**), a real-world RL framework that couples Corrective Value Learning (CVL) with Adjoint-based Policy Optimization (APO). Specifically, CVL exploits human corrections both to learn action values from their observed outcomes and to provide comparative supervision against rejected policy proposals. Guided by the learned critic, APO propagates action-value gradients along the VLA's native flow trajectories and converts them into local policy updates, while preserving the original deployment sampler. Across six real-world manipulation tasks, CAPO achieves the highest mean autonomous success of 93.3% and the lowest intervention rate among the evaluated methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.