acceptodds
Under review as a conference paper at ICLR 2027

CRAFT: Constrained Rollout Allocation from Feedback for Vision Language Action Reinforcement Learning

Abstract

Online reinforcement fine-tuning of Vision Language Action (VLA) policies is often constrained by limited training budgets, with only a finite number of policy update rounds and rollout opportunities available for multi-task reinforcement learning. In this setting, every rollout must be assigned to a specific task, making the allocation of limited rollout opportunities across tasks a direct determinant of training efficiency and final performance. Existing methods typically employ uniform task sampling, which may repeatedly sample already saturated tasks without producing further performance gains, leave difficult tasks generating homogeneous failures that provide little effective within-group learning signal, and fail to prevent previously acquired capabilities from deteriorating during subsequent training. To address these limitations, we propose CRAFT, a rollout allocation framework for budget-constrained Vision Language Action reinforcement learning. CRAFT uses feedback from completed rollout groups to adaptively distribute future rollout opportunities across tasks. Specifically, CRAFT jointly models target success deficits, within-group outcome variation, and differences between fast and slow success estimates to identify tasks with greater current training value. It further employs priority smoothing, coverage bonuses, and a minimum sampling probability to convert group feedback into a stable next-task distribution while preventing individual tasks from being persistently overlooked. Under identical training settings and matched evaluation protocols, CRAFT achieves consistent aggregate gains over uniform task allocation within fixed update budgets: after 30 updates, it improves LIBERO-Long success by 4.0%; after 10 updates, it improves LIBERO-Spatial success by 1.4%; after 30 updates, it improves the four-suite LIBERO average by 0.9%; and after 30 updates, it improves weighted LIBERO-PRO success by 7.6%. These results demonstrate that CRAFT allocates rollout opportunities more effectively for update-constrained multi-task reinforcement learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.