GRIP: A Generalist In-Context Reward Model for Robotic Manipulation
Abstract
Dense rewards provide step-level feedback for assessing progress, selecting demonstrations, and training policies, yet specifying them separately for every task is costly. We present GRIP, an in-context reward model that predicts a task's dense reward from a small set of labeled observations supplied at inference time. Its main configuration combines a frozen visual encoder with a frozen in-context regressor, allowing tasks to change through context rather than parameter updates. On MetaWorld and ManiSkill, GRIP improves dense-reward prediction over regressors fitted on the same examples. Its predictions improve with more context, depend on the observation–reward pairing, and remain stable when a shared context contains up to 50 tasks. On ProcVQA, a lightweight adapter trained on global visual features raises VOC from 0.6287 to 0.8133. GRIP's predicted rewards also select demonstrations that improve behavior-cloning success. Together, these results show how context-specified rewards can support progress assessment and data selection across robotic manipulation and embodied procedural tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.