acceptodds
Under review as a conference paper at ICLR 2027

OVOGO: Offline Value-Guided On-Policy Generative Optimization for Flow-Based Vision-Language-Action Models

Abstract

Flow-based vision-language-action models can benefit from robot experience collected during deployment, but obtaining new interactions for every policy update is costly. Offline reinforcement learning enables repeated reuse of collected experience, yet applying it to generative VLAs requires a critic that can evaluate newly generated actions and an effective way to improve the generative policy from these value signals. We propose OVOGO, an offline reinforcement learning framework that decouples offline value learning from on-policy generative optimization. OVOGO introduces a Return-Anchored Ranking Critic that combines in-sample distributional value learning with guarded local ranking, keeping values grounded in observed robot experience while improving discrimination among generated actions. The frozen critic then provides group-relative value signals for selected-transition optimization of fresh actions sampled from the current flow policy. Experiments on RoboTwin 2.0 and DexJoCo show that OVOGO consistently improves the pretrained VLA and outperforms existing offline adaptation and value-guided methods. The results suggest that expensive robot experience can be repeatedly reused for value learning while inexpensive generative sampling drives policy optimization, providing an effective approach to experience-based post-training of large VLAs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.