MENTOR-ISP: Stage-Wise Credit Mentoring for Single-Agent ISP Tuning
Abstract
An image signal processor (ISP) is a sequential pipeline in which upstream decisions affect downstream modules. Single-agent reinforcement learning efficiently optimizes ISP parameters jointly, but typically relies on a global reward that provides coarse credit across coupled controls. Serialized multi-agent approaches provide finer stage-wise supervision, but require repeated ISP execution at deployment. We present Mentor-ISP, a modular single-agent RL framework that decouples structured credit assignment and inter-stage dependency modeling from serialized ISP execution. Mentor-ISP factorizes the joint action into pipeline-ordered groups, uses sequential action handoffs to condition downstream decisions on upstream predictions, and applies stage-wise credit during training. At inference, all decisions are assembled into a single joint action, requiring only one ISP rendering per optimization step. Across two real-camera image-quality datasets, Mentor-ISP substantially improves peak signal-to-noise ratio (PSNR) over a matched single-agent baseline while approaching serialized multi-agent performance at near-single-agent inference cost. Ablations show that stage-wise credit provides the primary gain, with action handoff offering complementary benefit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.