Frequency Boundaries of Visual Semantics for Grounded Answering in Frozen Vision-Language Models
Abstract
Vision-language models answer questions across diverse visual tasks, yet how image frequency bands affect answer correctness and semantic content remains unclear. We introduce a paired analysis framework based on independent band suppression to characterize the frequency boundaries of visual semantics in frozen models. Holding the model and question fixed, we independently generate eight radial band-stop variants of each image and compare original and filtered responses using five semantic labels based on the question and reference answers, distinguishing preservation, degradation, and error correction. Across four tasks and five models from two families, degradation follows task-dependent frequency-band patterns shared across models. Some initially incorrect answers also become correct under band suppression. For shared instances answered correctly by all five models on unfiltered images, InternVL3 has significantly lower correct-to-incorrect transition rates than Qwen2.5-VL across multiple bands in text-centric tasks. Within families, larger models show similar reductions, mainly in document understanding. Human validation yields 91.72% weighted response-level agreement between final automatic labels and human judgments. These findings link suppressed bands to the direction and type of semantic transition, revealing how frequency boundaries vary with task, family, and scale.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.