acceptodds
Under review as a conference paper at ICLR 2027

HIVE-Bench: Evaluating Patch-Level Visual Representations for Egocentric Robot Manipulation

Abstract

Pretrained visual encoders are widely reused in robot manipulation, yet most evaluations read them through a pooled embedding of third-person images, whereas bimanual policies consume dense patch tokens from onboard cameras. We introduce HIVE-Bench, to our knowledge the first benchmark that compares dense patch-level representations for egocentric bimanual manipulation under one shared policy. A flow-matching action head reads the patch tokens of 20 encoders on 24 RoboCasa-GR1 and RoboTwin 2.0 tasks, across frozen and fine-tuned encoders, training regimes, and data, model, and layer sweeps. We further relate more than fifty diagnostics and published vision-benchmark scores to closed-loop success. Three findings run against common practice. Averaging each view's patch tokens into a single token removes most of the success of frozen DINOv2 and DINOv3. Measures of world reconstruction, from segmentation, depth, and correspondence scores to probes that decode robot state or object position, either do not track success or lose this relationship once vision-language models are included. In contrast, an action readout that decodes commanded motion from two head-camera frames tracks success in every setting we test, with Spearman absolute rho of at least 0.54. Finally, better vision is not necessarily better control: larger models and newer pretraining improve vision benchmark scores without consistently improving manipulation success.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.