acceptodds
Under review as a conference paper at ICLR 2027

Omni-Mono: Mutually Informed Reasoning across Vision, Language, and Action

Abstract

Vision-language-action (VLA) models increasingly use intermediate reasoning to guide robotic manipulation. However, existing approaches often restrict reasoning to one modality or a predefined sequence, limiting mutual feedback across modalities. We propose Omni-Mono, a framework combining joint multimodal reasoning with modality-specific refinement. The Omni Reasoner jointly generates textual, visual, and action representations in a shared latent space through grouped attention with modality relation masks, enabling information exchange across reasoning streams. The Mono Reasoner then refines these representations through dedicated experts, adaptively fusing guidance from Omni representations and preceding experts according to the current reasoning state. Together with the task context, the refined representations condition a flow-matching head to generate action chunks. Experiments on LIBERO, LIBERO-Plus, SimplerEnv, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved performance and robustness to environmental perturbations. Ablations and representation analyses further support the complementary benefits of joint reasoning and modality-specific refinement. The code repository is available at https://anonymous.4open.science/r/omni-mono/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.