DRAVS: Directional Visual Reward Learning with Adaptive View Selection
Abstract
Vision-language models provide visual supervision for reinforcement learning, reducing reliance on manually designed rewards. Multi-view videos further alleviate occlusion by revealing task-relevant motion and contacts that may be hidden from a single viewpoint. Existing video-supervised state-wise reward formulations directly add learned state scores to task rewards. However, policies trained with such feedback can exhibit late-stage performance degradation, losing part of their earlier gains. Because the immediate visual term depends only on the current state, it does not explicitly distinguish progress from regression across transitions. Moreover, uniform random sampling of supervision views treats cameras equally despite differences in the visibility of task-relevant behavior. To address these limitations, we propose Directional Visual Reward Learning with Adaptive View Selection (DRAVS). First, Direction-Aware Visual Reward replaces state-wise addition with consecutive-state score differences computed using a scorer held fixed within each reinforcement-learning phase, providing direct feedback on progress and regression without additional vision-language model queries. Second, Budgeted Relative View Selection balances exploitation and exploration through sparse cross-view comparisons to adaptively select supervision views, with at most 5% additional scored clips over one-view-per-segment supervision. Compared with the state-of-the-art Multi-view Video Reward method, DRAVS raises mean final success rate on the ten-task Continual World benchmark built on MetaWorld from 0.54 to 0.69 and mean final return across nine HumanoidBench tasks from 616.9 to 836.3. Learning curves further show that DRAVS mitigates late-stage performance degradation and better preserves earlier gains. Additional qualitative visual comparisons are available at the anonymous project page: https://anon-visual-research.github.io/DRAVS/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.