acceptodds
Under review as a conference paper at ICLR 2027

GOT-Align: Learning Generative Operation Tokenizers for Preference-Aligned Image Enhancement

Abstract

Photo retouching is inherently subjective: the same photograph can admit multiple plausible renditions, while different users may prefer different appearances. Existing methods either learn toward selected expert retouches or rely on explicit controls and instructions, making it difficult to capture implicit and spatially varying preferences. We propose GOT-Align, a preference-adaptive generative photo retouching framework that models retouching transformations in a shared spatial operation space. Specifically, a Generative Operation Tokenizer (GOT) discretizes bilateral affine transformations into reusable local operation tokens, whose spatial compositions represent diverse retouching transformations while preserving the underlying image content. A conditional autoregressive Transformer then models the image-conditioned distribution over operation-token maps to generate multiple plausible candidates for the same input. For preference adaptation, human or vision-language-model judges compare candidates over local image regions, preserving regional preference evidence that would otherwise be collapsed into a single image-level judgment. These regional preferences are attributed to spatially associated operation tokens and used for GRPO-based policy adaptation. Once adapted to a preference source, the model directly generates preference-aligned retouches for new inputs without requiring per-image controls or textual instructions. Experiments on standard photo retouching benchmarks demonstrate diverse generation and effective adaptation to different preference sources.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.