Learning to Train Underwater Image Enhancers with Reinforcement-Guided Update Scheduling
Abstract
Underwater images often suffer from color distortion and scattering-related detail degradation, which arise from different physical processes and exhibit distinct optimization dynamics. Existing methods that jointly optimize color correction and detail restoration typically rely on fixed loss weights or predefined update schedules, implicitly treating their relative importance as constant throughout training. However, the two objectives often progress asynchronously, making static coordination unable to accommodate their stage-dependent optimization demands. We propose a learning-to-train framework that formulates update scheduling as a Markov decision process. A PPO-based controller observes the current training state and dynamically selects among alternating, joint, and weighted updates for a cascaded enhancement network. We further construct a frequency-aware state representation from recent reconstruction quality trends and low/high frequency alignment to track the relative progress of global color recovery and local detail restoration. By maximizing cumulative rewards that account for reconstruction improvement, loss reduction, and frequency balance, the controller learns stage-adaptive update policies. Comprehensive qualitative and quantitative evaluations on multiple benchmark datasets, together with extensive ablation studies, demonstrate the effectiveness and robustness of the proposed method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.