CRAFT: Teaching Vision-Language Reasoners What Makes Reasoning Worth Keeping
Abstract
Multimodal reasoning models leverage Chain-of-Thought (CoT) reasoning to solve complex visual tasks, yet incorrect deductions, visually ungrounded descriptions, and irrelevant visual attention can degrade reasoning quality and incur unnecessary inference cost. We introduce CRAFT (Correctness, gRounding, And Focus Training), a discrimination-like auxiliary verification framework for multimodal CoT compression. Unlike existing approaches that directly construct, imitate, or optimize shorter reasoning chains, CRAFT shifts the training objective toward reasoning quality assessment. Specifically, it introduces three simple yet complementary auxiliary verification tasks targeting reasoning correctness, visual faithfulness, and visual focus. These tasks supervise the model to identify incorrect deductions, visually unsupported descriptions, and task-irrelevant visual information, thereby strengthening its internal ability to distinguish high-quality reasoning from flawed or redundant content. Through joint training, the learned verification capability transfers to open-ended multimodal reasoning, enabling the model to reduce redundant reasoning while preserving critical visual evidence, without requiring carefully constructed concise CoTs, reward models, or reinforcement learning. Experiments on four public multimodal question-answering benchmarks across multiple model scales show that CRAFT substantially reduces output token usage while maintaining answer accuracy, and further improves accuracy in most settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.