Generation as Evidence: Improving Visual Understanding in Unified Multimodal Models
Abstract
Unified multimodal models support visual understanding and image generation, but whether generation provides useful evidence beyond strong understandingonly methods remains unclear. We study this question through caption discrimination: selecting the correct description of an image from two candidates. We score the observed image conditioned on each caption using the generation pathway, then combine generation and understanding scores with a linear combiner calibrated on a small labeled set. The pretrained model remains frozen, and no new images are generated. We compare against four progressively stronger understanding-only baselines, with the strongest using the same understanding scores, calibration data, and combiner. Across six unified models and a crossmodel control on five compositional benchmarks, generation significantly improves 10 of 27 evaluated pairs by 1.2–8.1 percentage points over this baseline. Eight gains replicate in an independent implementation; uninformative replacements of generation scores yield no significant gains. The median gain falls from 4.8 to 1.1 points as understanding baselines are strengthened, showing why generation’s contribution must be measured against strong controls. Localization, occlusion, and image-shuffling analyses show that generation scores depend on image content and concentrate on regions that distinguish the captions. These findings show that generation pathways can provide complementary visual evidence for caption discrimination within frozen unified models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.