CAPTURE with PhotoEncoder: Photographic Understanding with Representation Enhanced Vision-Language Models
Abstract
Vision-language models (VLMs) could be powerful tools for understanding and retouching photographs, yet their ability to perceive fine-grained photographic attributes, such as brightness and color temperature, remains limited. Using the **Photographic Diagnostic Probe**, we reveal that existing VLM vision encoders inadequately capture these photographic attributes. To address this limitation, we introduce **PhotoEncoder**, which structures visual representations around photographic attributes and their controlled transitions while preserving pretrained semantic representations. We further construct a dataset of 5M vision-language pairs and 81K instruction-tuning examples derived from editing records to support training. By integrating PhotoEncoder into the language model and applying instruction tuning, we build **CAPTURE** for photographic perception, aesthetic understanding, and retouching assistance. Experiments show that CAPTURE better distinguishes subtle tonal and color changes, enabling more reliable aesthetic judgments and editing decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.