acceptodds
Under review as a conference paper at ICLR 2027

DRIVE: Dense Reward-Informed Action-Value Learning from Mixed Experience for Flow-Based VLA Policies

Abstract

Vision-Language-Action (VLA) policies are predominantly trained through imitation learning on demonstration data, while the rich execution signals contained in autonomous failures and human interventions remain underexploited. Such mixed execution experience contains both effective and low-value actions, making direct imitation unsuitable for policy refinement. We propose DRIVE, a dense reward-informed action-value learning framework for refining flow-based VLA policies from mixed robot experience. DRIVE first adapts a task-conditioned visual-language progress estimator on successful trajectories, then uses it to quantify advancement, stagnation, and regression across expert demonstrations, autonomous successes, autonomous failures, and successful human-intervention trajectories. The resulting dense supervision is distilled into an action-chunk critic that estimates long-horizon values for candidate action chunks. Under the same task condition, DRIVE compares candidate values and converts their relative differences into soft refinement labels. These labels guide a negative-aware forward flow-matching objective that reinforces high-value action chunks while explicitly suppressing inferior alternatives. Importantly, the progress estimator and critic are used only during training, so the refined policy incurs no additional inference-time computation. Experiments on LIBERO and four real-world manipulation tasks show consistent improvements across multiple flow-based VLA backbones, achieving 98.25% average success on LIBERO and 62.5% on real-world tasks, with a 40.0-percentage-point improvement over the corresponding base policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.