From Complementarity to Integration Value: A Supply–Conversion Framework for Multi-Model Reasoning
Abstract
Recent advances in large language models have made multi-model integration an important approach to enhancing reasoning. Additional models can contribute through two broad modes: online integration, exemplified by inference-time collaboration, and offline transfer, exemplified by teacher-guided supervised fine-tuning (SFT) followed by standalone deployment. Existing methods improve partner complementarity or teacher–student compatibility within each mode. Beyond these within-mode advances, there remains limited understanding of how the same candidate models create different integration value across roles and what governs these differences. To this end, we introduce a supply–conversion framework that evaluates shared candidate pools across online and offline modes against their corresponding single-model alternatives. For online collaboration, we quantify complementary coverage and selection loss to derive an exact condition for outperforming the stronger candidate. For offline SFT, we measure realized student improvement and the additional value of combining teachers relative to the better single-teacher alternative. Across diverse reasoning tasks, we find that collaboration gains depend jointly on complementarity and selection quality, teacher value follows a different structure from inference-partner value, and the better single teacher provides a strong reference for uniform teacher mixtures. Building on this analysis, we combine task-local probes with cross-pair judge calibration to select between single-candidate inference and collaboration, achieving lower selection regret than the evaluated score-, size-, and diversity-based baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.