acceptodds
Under review as a conference paper at ICLR 2027

FlexSense: Flexible Sensor Configurations with a Unified Vision-Language-Action Model

Abstract

Vision-Language-Action (VLA) models increasingly generalize across tasks, environments, and robot embodiments, yet most remain coupled to a fixed sensory interface. In practice, robots expose different combinations of optional sensors, and training a separate policy for every configuration scales poorly and prevents demonstrations collected under different sensing setups from being fully shared. We introduce FlexSense, a unified VLA framework for robotic manipulation under heterogeneous sensor configurations. It consists of three modules: First, Shared-Residual Sensory Encoding (ShareRes) captures transferable control structure through a shared sensory backbone while preserving sensor-specific cues through lightweight residual adapters gated by zero-initialized scalars. Second, Sensor-Set-Conditioned FiLM (SetFiLM) explicitly conditions each action block on the active sensor set, complementing token masking with configuration-aware action generation. Third, Utility-Guided Sensor Selection (UtilSelect) derives subset-dependent supervision from counterfactual action errors of the frozen policy, distills it into a lightweight scorer, and selects the sensor subset with the lowest predicted action error at each action-chunk query. We evaluate FlexSense on LIBERO, ManiSkill, and real-world robotic manipulation across multiple sensor configurations. Across LIBERO and ManiSkill, FlexSense remains within 1.6 percentage points of configuration-specific specialists in configuration-level average success rate across all tested sensor configurations. Real-robot experiments further show that the same policy benefits from complementary depth and tactile sensing, reaching 100% success on Pick-and-Place and 80% on Precision Plug Insertion. Our code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.