acceptodds
Under review as a conference paper at ICLR 2027

Frequency Boundaries of Visual Semantics for Grounded Answering in Frozen Vision-Language Models

Abstract

Vision-language models answer questions across diverse visual tasks, yet how image frequency bands affect answer correctness and semantic content remains unclear. We introduce a paired analysis framework based on independent band suppression to characterize the frequency boundaries of visual semantics in frozen models. Holding the model and question fixed, we independently generate eight radial band-stop variants of each image and compare original and filtered responses using five semantic labels based on the question and reference answers, distinguishing preservation, degradation, and error correction. Across four tasks and five models from two families, degradation follows task-dependent frequency-band patterns shared across models. Some initially incorrect answers also become correct under band suppression. For shared instances answered correctly by all five models on unfiltered images, InternVL3 has significantly lower correct-to-incorrect transition rates than Qwen2.5-VL across multiple bands in text-centric tasks. Within families, larger models show similar reductions, mainly in document understanding. Human validation yields 91.72% weighted response-level agreement between final automatic labels and human judgments. These findings link suppressed bands to the direction and type of semantic transition, revealing how frequency boundaries vary with task, family, and scale.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.