acceptodds
Under review as a conference paper at ICLR 2027

Beyond Relevance: Rethinking What Robots Need to See Across Tasks and States

Abstract

Visual relevance does not determine observation precision or visual capacity. Motivated by action sensitivity to state errors, we study action-conditioned measurement sufficiency beyond the Base policy’s existing context. Controlled ACT analyses show task- and state-dependent location–precision allocation under a fixed budget, no uniformly monotonic benefit from increasing the precision ceiling or residual capacity, and different preferred configurations across tasks. We propose Action–Visual Precision Matching: action context jointly ranks location–precision candidates, Global Top-K selects them under a shared budget, and only selected high-precision observations are encoded. Their feature residuals relative to base-aligned references refine native action generation. The main ACT configuration improves mean success rate by 2.65 percentage points over its matched Base and outperforms same-budget random allocation. A post-hoc task-wise configuration combination reaches 84.20%, exceeding Base and the best uniform configuration by 4.03 and 1.38 percentage points. Diffusion Policy and π0 further test integration across action generation structures. These results support matching visual resources to current action requirements rather than uniformly maximizing them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.