acceptodds
Under review as a conference paper at ICLR 2027

ChainDiff: Controllable Fine-Grained Image-Pair Synthesis and Frequency-Domain Difference Attention for Image Difference Captioning

Abstract

Fine-grained Image Difference Captioning (IDC) requires precisely describing sub-part-level differences between two highly similar images. Existing work is constrained by two bottlenecks: fine-grained image pairs are scarce and their differences are uncontrollable, and mainstream Vision-Language Models (VLMs) rely on implicit alignment to discover weak difference signals hidden inside overwhelming commonality. We propose ChainDiff, a controllable image-pair synthesis pipeline that combines dual-channel similarity filtering, a hard upper bound on the changed-region ratio, and chained editing with segment sampling to make sub-part-level fine-grainedness and per-pair change counts explicitly controllable. We further propose the Frequency-Domain Difference Attention (FDDA) module, which decouples multi-scale difference signals via a feature-domain discrete cosine transform (DCT) decomposition and models differences explicitly via dual-branch image-pair cross-attention, integrated into VLMs in a plug-in manner during supervised fine-tuning (SFT). We term the resulting models, trained on ChainDiff data with FDDA integrated, HarmonicDiff. Across three base VLMs and eight test sets, HarmonicDiff attains the best CIDEr among Multimodal LLM (MLLM) methods on five test sets; specialized non-MLLM baselines remain stronger in CIDEr only on synthetic CLEVR-Change.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.