C-VOIR: Task Inference for Heterogeneous Robot Teams
Abstract
Behavior foundation models (BFMs) provide reusable policies, but an unknown deployment reward requires evidence before a policy can be selected. We study online task inference in heterogeneous robot teams whose members have different embodiments, sensors, actions, and probe costs, and communicate only with neighbors. We propose C-VOIR, a cost-aware value-of-information method that weights reports by sensor noise, age, and prediction consistency. It values physical probes by their expected reduction of uncertainty between competing decision labels minus physical cost. Robots exchange their best positive bids, and a strict-majority lease authorizes at most one probe. The team executes a common policy label only when fresh, consistent opinions indicate that no worthwhile probe remains and the decision gap supports a winner; otherwise, it retains baseline actions and no-op probes. Across three evaluation trials, C-VOIR achieves the highest reported epoch-10 running-average return on four open-sourced benchmarks. We also introduce HeteroArm-Mission, a 600-instance MuJoCo benchmark with two- and three-joint arms, heterogeneous observations, and physical probes. C-VOIR achieves the highest reported non-Oracle task return 0.8810 and cost-adjusted return 0.8750, while reducing mean probe cost by relative to Information Gain. Our code is open-sourced at https://anonymous.4open.science/r/Anonymous_code-77E6.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.