acceptodds
Under review as a conference paper at ICLR 2027

VEFR: Visual Evidence Fusion Routing for Budgeted Skill Selection in Fine-Grained Visual Recognition with MLLMs

Abstract

Fine-grained recognition requires distinguishing subtle visual differences, an MLLM-based system can approach the same image through several recognition strategies. These strategies differ in cost and have complementary strengths across classes and individual images. We formulate recognition as budgeted skill routing: selecting an executable recognition strategy before observing its output or incurring its execution cost. We propose Visual Evidence Fusion Routing (VEFR), which estimates skill competence from two sources of visual evidence. Class-conditioned statistics capture recurring skill preferences, while correctness records from visually similar profiling images provide query-specific evidence. A soft visual class distribution weights the class statistics, and a validation-selected fixed weight combines the two estimates without training a separate routing network. We separate competence estimation from deployment control, allowing different prioritization of accuracy–cost without recalculating the evidence. Across eight image recognition benchmarks, VEFR achieves better accuracy compared with fixed skill and skill routing methods while consuming much fewer MLLM tokens. These results show that visual evidence can guide skill selection to improve the accuracy–cost trade-off.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.