Dual Tuning: Joint Screening of Data Utility and Supervision Modes for Multimodal Fine-Tuning
Abstract
Multimodal fine-tuning requires decisions about both the training data and the form of supervision, such as chain-of-thought (CoT) or direct-answer (DA) supervision. These decisions are jointly shaped by the capability of the base model, the quality of CoT traces, and the characteristics of the target task, rather than by any fixed recipe. Consequently, isolated ablations insufficient for joint assessment. Here we introduce Dual Tuning, a screening protocol conditioned on the training configuration that addresses a complementary question: for a given base model and candidate dataset, which task groups benefit from training, and which supervision mode is better supported by the observed gains? It uses one joint supervised fine-tuning run per configuration to obtain gains over the base model and differences between CoT and DA modes for each task group. Experiments on spatial, mathematical, and multidisciplinary tasks reveal heterogeneous patterns across subgroups and models. Qwen2.5-VL-7B favors DA under the evaluated spatial configuration and CoT on MathVista. However, a CoT advantage over DA does not always coincide with a genuine improvement over the base model. The results on the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark identify 9 subjects favoring CoT, 4 favoring direct answers, and 10 with negative gains in both modes. Finally, subset retraining experiments provide directional support for using these signals in data curation. Dual Tuning offers a common empirical basis for screening candidate data and assigning supervision modes before subsequent training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.