One Object, Different Corrections: Controllable Low-Light Enhancement with Photometric Actions
Abstract
A vision-language model (VLM) brings semantic understanding to low-light image enhancement, but language guidance alone does not specify how much correction to apply. An instruction to brighten an object leaves open how its shadowed and well-lit parts should be adjusted. We introduce an Executable Photometric Action Space (EPAS) that separates image-dependent correction from semantic execution control. A signed log-gain field specifies correction direction and magnitude, while a separate control field determines where and how strongly it is applied, allowing spatially varying corrections within semantic regions. During training, Metropolis–Hastings sampling explores regional actions under an energy defined on the input image's photometric statistics and VLM-derived preferences, with VLM quality assessment guiding candidate selection. This constructs action targets without paired normal-light references; the selected actions are distilled into a deterministic predictor for inference without action search. Applying the controlled gain produces a preview that guides generative restoration for detail recovery and noise suppression. Experiments show that spatially varying gains reduce the shadow undercorrection and highlight overcorrection caused by object-wise uniform gains, especially when correction requirements vary substantially within an object. Outputs approximately follow supplied actions, with more accurate restoration when using the input’s own gain field. A single trained model achieves competitive reconstruction quality, strong perceptual performance, and flexible global and local control across diverse benchmarks and challenging illumination conditions without dataset-specific retraining or fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.