BARM: Bayesian Assistance Routing for Embodied Robot Manipulation over VLM and Human Experts
Abstract
Embodied agents deployed in open-world environments must decide, for each incoming task, whether to execute it autonomously with a lightweight policy, invoke a vision-language model planner, or request human assistance. These execution modes differ in capability and cost, and the routing controller is typically trained on atomic tasks but deployed on compositional, multi-step instructions that lie outside its training distribution. We cast embodied assistance routing as a contextual bandit problem with heterogeneous arm costs and identify two desiderata for sample-efficient routing under distribution shift: distance-aware uncertainty, so that unseen tasks are recognized as uncertain, and probability-matching exploration, so that uncertainty translates into cost-aware action. We propose BARM (Bayesian Assistance Routing for embodied robot Manipulation), a unified Bayesian framework whose single belief state simultaneously satisfies both desiderata: its epistemic variance grows with distance from the training distribution, and Thompson Sampling on this belief realizes probability-matching exploration over net rewards. Reward modeling, exploration, and online adaptation all emerge from this one belief rather than from separately engineered heuristics. We also establish a sublinear Bayesian regret bound for BARM under the well-specified regime (linear reward). On CALVIN and LIBERO benchmarks, BARM achieves the highest cumulative reward against 22 baselines spanning failure detection, reinforcement learning, and contextual bandit methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.