\(\pi\)-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
Abstract
Flow-based vision-language-action (VLA) models have emerged as a powerful paradigm for embodied control, but multi-step denoising makes policy-gradient online reinforcement learning (RL) difficult because the final-action likelihood induced by the sampler is not directly tractable. We identify an explorationsupervision mismatch within the denoising trajectory, where stochastic differential equation (SDE) sampling explores a wider local stochastic transition space while supervision derived from the final denoised output does not directly use the adjacent transitions observed during sampling. We propose π-StepNFT (Step-wise Negative-aware Fine-Tuning), a flow-native online RL framework that couples wider stochastic exploration with finer step-wise supervision without computing the final-action likelihood. π-StepNFT treats observed SDE transitions as supervision units, using variance-normalized transition errors to rank mirrored positive and negative velocity branches and map rollout feedback to local velocity updates. On LIBERO, starting from shared few-shot SFT initializations, π-StepNFT raises average success from 57.6% to 92.0% for πо and from 77.1% to 94.0% for πo.5. On ManiSkill, it achieves the highest reported mean success among the evaluated RL baselines on five test-only visual perturbations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.