Compact Chain-of-Thought Distillation with Visual Grounding Guidance for Vision-Language Models
Abstract
Multimodal chain-of-thought (CoT) reasoning has significantly improved the performance of large vision-language models (VLMs) in solving complex tasks. However, distilling this reasoning capability into small VLMs remains challenging. In this paper, we construct a large-scale multimodal CoT dataset comprising 120k samples, each featuring a compact CoT with a standardized structure, explicit length constraints, and region-level visual grounding annotations. Based on our dataset, we obtain two critical observations: (1) VLM layers vary substantially in their receptiveness to visual grounding knowledge; (2) Compact CoTs achieve greater reasoning efficiency and higher accuracy than lengthy native CoTs. Motivated by this, we propose the Visual Grounding-Guided CoT (VGG-CoT) distillation framework, which first identifies receptive layers via per-layer probing and restricts LoRA updates to these layers for efficient visual grounding transfer. Then, a standard LoRA distillation is performed over all layers to transfer multimodal reasoning capabilities. In this way, VGG-CoT decouples visual grounding from multimodal reasoning, tailoring the distillation process to the layer-wise receptiveness and limited capacity of small VLMs. Extensive experiments on several VLM architectures, including QwenVL and InternVL, demonstrate that VGG-CoT consistently outperforms native CoT distillation in terms of reasoning accuracy and token efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.