acceptodds
Under review as a conference paper at ICLR 2027

CAPTURE with PhotoEncoder: Photographic Understanding with Representation Enhanced Vision-Language Models

Abstract

Vision-language models (VLMs) could be powerful tools for understanding and retouching photographs, yet their ability to perceive fine-grained photographic attributes, such as brightness and color temperature, remains limited. Using the **Photographic Diagnostic Probe**, we reveal that existing VLM vision encoders inadequately capture these photographic attributes. To address this limitation, we introduce **PhotoEncoder**, which structures visual representations around photographic attributes and their controlled transitions while preserving pretrained semantic representations. We further construct a dataset of 5M vision-language pairs and 81K instruction-tuning examples derived from editing records to support training. By integrating PhotoEncoder into the language model and applying instruction tuning, we build **CAPTURE** for photographic perception, aesthetic understanding, and retouching assistance. Experiments show that CAPTURE better distinguishes subtle tonal and color changes, enabling more reliable aesthetic judgments and editing decisions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.