Twinsight: Guiding Human Teaching and Policy Adaptation with Learner Forecasts
Abstract
Successful demonstrations can still leave a robot policy's weaknesses unaddressed. Without seeing how the learner would continue, human teachers may miss useful corrections, and data selection cannot recover examples that were never demonstrated. We introduce Twinsight, a framework that connects human teaching and policy adaptation through learner forecasts. During human control, a learner-specific predictive model forecasts end-effector motion beyond the current action chunk, helping teachers anticipate later difficulties before committing to an approach. Persistent teaching bubbles record local differences between forecast and demonstrated motion and link them to coherent human action segments for subsequent selection and training. Across six simulated and three real-robot manipulation tasks, Twinsight achieves mean success rates of 74.0% and 80.7%, exceeding the strongest evaluated baselines by 5.3 and 6.7 percentage points, respectively. On shared simulation datasets, segment selection achieves similar mean success to training on all retained data with 30.7% fewer training steps.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.