acceptodds
Under review as a conference paper at ICLR 2027

SARM2: Multitask Stage Aware Reward Modeling for Self Improving Robotic Manipulation

Abstract

Fine-tuning vision-language-action (VLA) policies for long-horizon manipulation still largely depends on behavior cloning (BC), which requires costly high-quality demonstrations and limits policies to the demonstration distribution. Reward models support learning beyond demonstrations through data reweighting and dense reward signals for reinforcement learning (RL). However, existing reward models struggle to combine multitask scalability and precision: task-specific stage-aware models require per-task annotation and retraining, while general vision-language reward models often miss subtle progress in long-horizon tasks. We introduce SARM2, a multitask stage-aware reward model combining a shared action-primitive stage estimator with a multi-gate Mixture-of-Experts (MMoE) value head for accurate dense rewards. We use its dense rewards in SPIRAL, an offline-to-online pipeline for policy self-improvement from demonstrations and autonomous rollouts. On ten tasks, SARM2 reduces overall value-estimation MSE by 33.3% versus multitask SARM and 53.5% versus the best evaluated general-purpose reward model baseline. Starting from offline RL policies, SPIRAL raises success rates from 60% to 100% on Folding Flattened Shorts, 33.3% to 66.7% on Folding Crumpled Shorts, and 50% to 90.0% on Cleaning Whiteboard, supporting a robot data flywheel driven by accurate dense rewards.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.