acceptodds
Under review as a conference paper at ICLR 2027

Beyond Omni: A Unified MLLM for Embodied Sensory Understanding

Abstract

Recent embodied systems increasingly build on multimodal large language models (MLLMs) for their pretrained understanding and reasoning capabilities, but their perception of the physical world remains largely limited to basic modalities such as vision and audio. However, broadening MLLMs to more diverse sensory modalities from the physical environment is challenging: their representations are often weakly connected to the MLLM and vary substantially in structure and statistics, while their interpretation still requires shared high-level understanding and reasoning. We therefore propose UniSense, a unified embodied MLLM framework that supports diverse physical sensory modalities within a shared multimodal architecture. UniSense follows a two-stage training process: (i) aligning sensory representations with language to establish a language-accessible interface; and (ii) tuning modality-specific experts within the shared LLM to provide specialized processing for different sensory modalities. To support broad sensory learning, we curate UniSense–1M, a large-scale embodied sensory training corpus comprising 1M samples across 8 sensory modalities. We further evaluate UniSense on 12 benchmarks spanning depth, point cloud, IMU, tactile, thermal, radar, motion, and smell understanding. Across these heterogeneous sensing tasks, UniSense achieves strong performance and outperforms both generalist and specialized models on most benchmarks, demonstrating the potential of pretrained MLLMs as unified models for embodied sensory understanding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.