Swimbird: Eliciting Switchable Reasoning Mode in Hybrid Autoregressive MLLMs
Abstract
Multimodal Large Language Models (MLLMs) have advanced perception and reasoning by bridging vision and language, yet most still rely on textual chain-of-thought (CoT), which is ill-suited to vision-intensive tasks. Recent methods inject a fixed number of continuous hidden states as visual thoughts to strengthen visual reasoning but often weakening text-based logical reasoning. We argue that the main limitation is a rigid, pre-defined reasoning pattern that cannot choose the most suitable thinking modality for each query. We introduce SwimBird, a reasoning-switchable MLLM that dynamically selects among three input-conditioned modes: (1) text-only reasoning, (2) vision-only reasoning with continuous visual thoughts, and (3) interleaved vision–text reasoning. SwimBird adopts a hybrid autoregressive formulation that unifies next-token prediction for textual thoughts with next-embedding prediction for visual thoughts, and uses a reasoning-mode curation strategy to construct SwimBird-SFT-92K, a diverse SFT dataset covering all three reasoning patterns. By enabling query-adaptive mode selection and adaptive latent-token allocation, SwimBird preserves textual logic while boosting vision- dense reasoning. Experiments across textual reasoning and challenging visual understanding benchmarks show that SwimBird achieves state-of-the-art results and robust gains over prior fixed-pattern reasoning methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.