Arc-Smoothed Caption Supervision for Fine-Grained Spatial CLIP
Abstract
CLIP is the image encoder behind zero-shot recognition and many multimodal large language models (MLLMs), yet it is weak at fine-grained spatial perception such as counting objects and judging their positions and relations. Therefore, classifiers and MLLMs that reuse its encoders inherit this weakness. Existing remedies fine-tune the image encoder against a second large model placed in the training loop, such as a frozen diffusion model or a vision-only teacher, and give the text encoder no spatial signal. We propose arc-smoothed caption supervision, a simple alternative whose training targets are all points in CLIP's own image–text embedding space and which uses no other network during training or inference. For each training image, we generate a spatial caption once, offline. During fine-tuning, we move along the great-circle arc between the embeddings of the class template and the spatial caption to a position drawn from a U-shaped Beta law, and use the resulting point as the positive target of CLIP's standard contrastive loss. Targets inside the arc are essential: using the same two captions as separate positives at the same exposure gives lower spatial zero-shot accuracy than template-only fine-tuning. We show that the mixing law trades fidelity to the two caption objectives against an interior term that contrastively supervises the midpoint between each image's template and its own caption, and controlled comparisons are consistent with both sides of this trade-off. On CLIP ViT-L/14, relative to matched template-only fine-tuning and across two captioners, arc smoothing improves spatial zero-shot accuracy by – points and text-to-image retrieval R@1 by – points, while keeping zero-shot transfer on datasets at the template-only level. Used as the visual encoder of LLaVA-1.5, it improves average VQA accuracy by points, including points on TallyQA counting, suggesting that spatial information supplied through language during encoder fine-tuning carries over to downstream multimodal reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.