acceptodds
Under review as a conference paper at ICLR 2027

From Visual Facts to Semantic Concepts: Structuring Frozen Vision–Language Representations with Evidence–Concept Decomposition

Abstract

Frozen vision–language models provide transferable visual features that encode both directly observable cues and higher-level semantic information. We organize one frozen representation into shared, evidence, and concept pathways while keeping the pretrained backbone fixed. Raw visual fact descriptions are mapped to typed evidence tags covering entities, attributes, actions, relations, compositions, and scenes, while abstract concept labels are normalized separately. The model uses specialized residual branches, a concept-conditioned evidence gate, prototype-based predictors, and evidence-aware fusion. On a Flickr30k-based fact–concept benchmark, the full model reaches competitive balanced mAP (19.68 versus 19.29 for a strong shared MLP) while providing separate pathways that can be removed, perturbed, swapped, and probed. A fact bottleneck retains strong fact performance but drops to 7.41 concept mAP. Source removal, gate perturbation, evidence-type analysis, and probing show distinct roles for shared, evidence, and concept information, with the clearest gate effect appearing in the evidence-to-concept route. The structured pathways retain competitive prediction while making evidence and concept use directly measurable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.