acceptodds
Under review as a conference paper at ICLR 2027

RoboStressBench: A Diagnostic Benchmark for Visual Stress in Embodied Scene Understanding

Abstract

Vision-language models (VLMs) used in embodied systems must interpret scenes under challenging visual conditions, yet it remains difficult to distinguish failures caused by visual stress from those caused by the underlying task or scene. We introduce RoboStressBench, a diagnostic benchmark for evaluating VLM robustness to visually challenging conditions in robot-centric scenes. An image-formation-inspired taxonomy organizes visual stress into four dimensions—Material, Viewpoint, Lighting, and Geometry—and 16 fine-grained categories, without assuming that these factors are independent in natural images. RoboStressBench contains more than 7K examples collected through naturally occurring stress cases, targeted stress synthesis, and additional real-world collection, and evaluates complementary capabilities including target and placement grounding, spatial reasoning, state understanding, and action-selection-style question answering. To complement broad stress coverage, we further evaluate matched nominal–stressed scene-question pairs, enabling direct measurement of performance degradation under controlled visual modifications. Evaluating 16 VLMs reveals substantial variation across both stress types and task formats, with different visual conditions exposing distinct weaknesses that are obscured by aggregate scores. We additionally study StressDART, a test-time rectification baseline that processes visually stressed inputs before reasoning, and analyze both its performance gains and the risk of altering task-relevant evidence. RoboStressBench targets the visual components relevant to embodied systems, rather than end-to-end closed-loop robot reliability, and provides a structured testbed for studying and improving multimodal robustness under challenging visual observations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.