CATD-MoE: Conflict-Aware Timestep-Dynamic K/V Mixture-of-Experts for Instruction-guided Image Editing
Abstract
Instruction-guided image editing models based on diffusion Transformers generally face a trade-off between editing flexibility and source image fidelity. Existing approaches typically rely on fixed K/V injection or manually defined task labels, making them less capable of adapting to dynamic editing requirements across different denoising stages and network depths. To address this issue, we propose CATD-MoE, a Conflict-Aware Timestep-Dynamic K/V Mixture-of-Experts framework that requires no retraining of the diffusion backbone. Specifically, we model the source image K/V and target denoising-state K/V as two memory experts, while their continuous interpolation forms a controlled bridging region. A lightweight Router dynamically determines their relative contributions based on denoising progress, mask properties, K/V conflict, and text representations. All diffusion backbone parameters remain frozen, with only the routing network being trained. During training, counterfactual candidates are constructed offline using target images, and Pareto-optimal supervision is derived according to editing quality, source image fidelity, and boundary quality, enabling the model to learn dynamic K/V mixing ratios. Experimental results on the independent PIE-Bench benchmark show that CATD-MoE effectively improves non-edited region fidelity and reduces boundary errors while maintaining high semantic consistency with the editing instructions. Further fine-grained evaluations and ablation studies demonstrate that joint spatiotemporal adaptation and K/V conflict awareness are key contributors to the observed performance gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.