Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
Abstract
Ensemble of either large language models or large vision models has been studied extensively to date. However, there is an increasing demand on tackling the problem of ensemble of diverse Vision-Language Models (VLMs) based on both vision and language modalities jointly. This paper presents V3Fusion, a Vision-Verification enhanced VLM Fusion framework for high performance visual reasoning with three unique contributions. First, V3Fusion introduces focal error diversity to capture complementary reasoning across VLMs by combining language-based focal error diversity with a Focal Central Kernel Alignment (CKA) based focal diversity metric (CKA-focal) to measure disagreement in visual embeddings. Second, on the constructed ensemble surface from a pool of candidate VLMs, we applied a Genetic Algorithm to effectively prune out those component VLMs that do not add value to the fusion performance. Finally, we identify the best combination for each reasoning task and resolve conflict among the outputs of multiple VLMs in the ensemble model pool, enabling heterogeneous models to capture epistemic uncertainty dynamically while mitigating hallucinations. Our V3Fusion approach is capable of producing dual focal-diversity fused predictions with high performance for vision-language reasoning, even when there is no majority consensus or the majority of VLMs make incorrect predictions. Extensive experiments with four popular VLM benchmarks (A-OKVQA, MMMU, MMMU-Pro, and OCR-VQA) validate the effectiveness of V3Fusion. The results show that V3Fusion outperforms the best-performing VLM on MMMU by 4.54% and MMMU-Pro by 2.29% gain in accuracy. For generative tasks, V3Fusion outperforms Intern-VL2-8b and Qwen2.5-VL-7b, the top-2 VLM performers on both A-OKVQA and OCR-VQA. Code is available at anonymous.4open.science/r/v3fusion-30D1/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.