acceptodds
Under review as a conference paper at ICLR 2027

\(\pi\)-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs

Abstract

Flow-based vision-language-action (VLA) models have emerged as a powerful paradigm for embodied control, but multi-step denoising makes policy-gradient online reinforcement learning (RL) difficult because the final-action likelihood induced by the sampler is not directly tractable. We identify an explorationsupervision mismatch within the denoising trajectory, where stochastic differential equation (SDE) sampling explores a wider local stochastic transition space while supervision derived from the final denoised output does not directly use the adjacent transitions observed during sampling. We propose π-StepNFT (Step-wise Negative-aware Fine-Tuning), a flow-native online RL framework that couples wider stochastic exploration with finer step-wise supervision without computing the final-action likelihood. π-StepNFT treats observed SDE transitions as supervision units, using variance-normalized transition errors to rank mirrored positive and negative velocity branches and map rollout feedback to local velocity updates. On LIBERO, starting from shared few-shot SFT initializations, π-StepNFT raises average success from 57.6% to 92.0% for πо and from 77.1% to 94.0% for πo.5. On ManiSkill, it achieves the highest reported mean success among the evaluated RL baselines on five test-only visual perturbations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.