LLaVA-MoR: Adaptive Visual Recursion for Multimodal Reasoning
Abstract
Multimodal Large Language Models (MLLMs) such as LLaVA have demonstrated promising capabilities in visual understanding and instruction following. Although allocating reasoning budgets according to problem difficulty has achieved better accuracy-efficiency trade-offs in text reasoning, whether and how to distribute adaptive budgets across visual tokens remains an open question in multimodal scenarios. In this work, we introduce LLaVA-MoR, the first approach to allocate question-conditioned recursion budgets to visual tokens in instruction-tuned MLLMs, optimizing multimodal reasoning by effectively preventing the model from overthinking simple visual regions and under-analyzing complex ones. To achieve this, we first propose the Question-Conditioned Visual Router, which explicitly models question–visual interactions to route visual tokens to different recursion depths. Furthermore, we propose Cosine-Guided Multi-Depth Adaptation, which gradually increases exposure to deeper recursion under answer supervision to adapt the backbone for visual processing while preserving its language capabilities. Extensive experiments demonstrate that LLaVA-MoR achieves a superior accuracy-efficiency trade-off across multiple backbones, even outperforming fixed maximum-depth recursion. These results highlight the critical value of token-level budget allocation in multimodal reasoning tasks. Code is available at \url{https://anonymous.4open.science/r/LLaVA_MoR/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.