Information-Guided Reward Learning for Contact-Rich Robotic Manipulation
Abstract
Contact-rich robotic manipulation requires distinguishing valid progress from ineffective motion. Yet correct insertion, misalignment, and jamming can appear visually similar, making image-based rewards unreliable. We propose an information-guided reward learning framework, which combines a physically supervised visual reward model with an action-history-based risk estimator. Privilege physical information, such as pose and contact signals, provide physical labels for reward learning, and do not appear during deployment. A dual-view visual model learns to predict positive, negative, or unclear physical progress, assigning unclear targets to visually ambiguous segments. The risk estimator learns whether visual positives are physically negative, based on executed-action history and visual class probabilities. At deployment, the risk estimator calibrate the visual reward model by changing high-risk positives to unclear. In Isaac Factory PegInsert, the framework suppresses 83.3% of negative-to-positive errors while retaining 82.9% of correct positives. Adding the learned reward to PPO increases aggregate success from 9.72% to 13.49%. On real Franka PegInsert data, Bayes-derived supervision increases mean episode-level student–teacher agreement from 43.70% to 59.32%, indicating improved supervision learnability rather than demonstrated gains in task success. These results support combining physical supervision and action history for more reliable contact-rich rewards.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.