acceptodds
Under review as a conference paper at ICLR 2027

Policy-Dependent View Acquisition

Abstract

Selecting camera frames before transmission requires a robot to estimate the value of observations it has not yet received. That value depends on both the views and the policy that consumes them. We formulate conditional view-set value and distinguish decision-level interaction from excess interaction induced by a policy. Selective substitutability training (SST) penalizes excess interaction together with subset regret and action-uncertainty mismatch; a residual coalition scheduler (RCS) selects fresh views from low-rate camera sketches under a per-decision byte cap. Across eight MuJoCo/robosuite manipulation tasks and three training seeds, SST+RCS achieves 69.58% episode success at a 55% communication cap, compared with 62.92% for all-subset training with a Set Transformer. The paired difference is 6.67 percentage points, with a 95% hierarchical bootstrap interval of [2.78,10.56]. A separate eight-seed evaluation shows a similar intermediate-budget gain. An expanded physical evaluation yields a +7.81-point paired difference across 576 trials per method. In the main frontier, SST+RCS loses at the smallest cap, ties at full view and has slightly higher mean latency than the reference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.