acceptodds
Under review as a conference paper at ICLR 2027

Same Scene, Different Decisions: Material Sensitivity in Embodied Models

Abstract

Embodied models must connect visual appearance to task-relevant decisions. We study their material sensitivity: how those decisions change when scene geometry is retained but materials are converted to a different representation. In 99 paired navigation episodes, success falls from 52.5% to 48.5%, yet 30 episodes change success status: gains and losses largely cancel in the average. We investigate this hidden turnover through paired visual grounding and navigation experiments. A readout study across six vision-language model configurations separates localization from response-format and coordinate-scale effects. Controlled two-view tests then require models to follow a queried object across positions, and target-window crossovers reveal that changes to the target and its surroundings can shift grounding in opposite directions. In navigation, 17 previously successful episodes fail while 13 previously failed episodes succeed under converted materials. These results put the identity of successful decisions alongside mean accuracy as a central object of evaluation. Material representation is therefore more than an asset-format choice: it is an input variable through which we can study the visual dependence of embodied models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.