HEPT: Human Execution Preference Transfer for Robot Policy Learning
Abstract
Human demonstrations teach robots more than task completion: they reveal repeat?able, source-associated differences in how tasks are executed. We study Human Execution Preference Transfer (HEPT) through a traceable interface from com?pleted teleoperation to a frozen robot policy. Without pairwise preference labels, language annotations, or operator identity as a model feature, a group-aware resid?ual model predicts speed and physical-time tool-center-point jerk from kinematic and contact summaries on held-out tasks. We align these coordinates with bounded interventions around a frozen vision-language-action policy and evaluate local ef?fects through exact-prefix counterfactual rollouts, separating outcome-aware reach?ability from deployable selection. In an independent RH20T confirmation with 168 demonstrations across 14 held-out tasks, the source representation achieves 0.6537 task-macro normalized error, compared with 0.7619 for an observable-state baseline, 0.8133 for a task mean, and 1.1136 for source swaps. Across five simula?tion task families, an outcome-aware counterfactual Oracle finds a task-admissible positive local response for both coordinates at all 90 tested anchors. On 20 fresh physical groups, a frozen matched source-profile prior yields utility of 0.4804 [0.1649, 0.8328] and exceeds source swap by 0.6202 [0.0889, 1.1880], with all matched selections locally task admissible. HEPT thus identifies source-specific execution information and demonstrates its task-preserving local realization; whole?episode policy adaptation remains a distinct problem.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.