DVCA: Denoising-Variance Credit Assignment for Flow-Based Robot Policy Optimization
Abstract
Recent Vision-Language-Action (VLA) models and World-Action Models (WAMs) use flow matching to model complex continuous action distributions. Online post-training can further improve these policies, but trajectory-level feedback offers limited guidance for prioritizing actions during training, limiting sample efficiency and performance gains. Process reward models and critics provide finer-grained feedback at the cost of additional training. To address this challenge, we introduce Denoising-Variance Credit Assignment (DVCA), which exploits the policy's denoising process for action-level credit assignment. Denoising variance, defined as the variance of clean-action estimates across denoising steps, serves as an intrinsic proxy for action uncertainty. Our analysis reveals that high denoising variance is concentrated in a small subset of actions associated with key phases of task execution. This observation motivates reweighting advantages or supervised losses to prioritize high-uncertainty actions. DVCA reuses intermediate predictions generated during online rollouts to estimate action uncertainty without additional training or inference. DVCA is a plug-and-play addition to existing post-training methods (e.g., GRPO and RL Token) across flow-based policies (e.g., and Fast-WAM). Simulation and real-world experiments demonstrate consistent gains in sample efficiency or task success rate over baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.