SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
Abstract
Recent works show that Vision-Language Models (VLMs) like CLIP struggle with logical reasoning, which can contribute to unreliable model behavior and hallucinations. An example of such logical operators is negation. Given a prompt like “retrieve (or generate) a street scene without pedestrians”, VLMs often fail to respect the “not”. Existing methods try to address the logical blind spots of VLMs by fine-tuning models on larger and richer datasets. However, this fine-tuning can compromise zero-shot generalization and often fails to fully eliminate these blind spots. In this paper, we use the negation operator as a case study and show that representing logical operators as traditional point embeddings in a joint embedding space can be limiting. Instead, we propose subspace modeling of negation, where text representation is a subspace in the embedding space rather than a traditional vector (point) embedding. Then we derive a closed-form Mahalanobis distance on a curved surface from an image embedding (point) to a text representation (subspace), replacing the conventional cosine similarity between image and text embeddings. Across retrieval and multiple-choice-question (MCQ) tasks, our new modeling improves negation understanding in VLMs including CLIP, SigLIP, and AIMV2 by about 30% on average over the baseline, without any additional re-training of the VLM, while preserving the model's zero-shot performance. More broadly, our framework could serve as a starting point for research on subspace modeling of other logical operators, not only at inference time, but also during model pre-training. Code is included in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.