acceptodds
Under review as a conference paper at ICLR 2027

CSTRA-OPD: Counterfactual Spatio-Temporal Alignment for Region-aware On-Policy Distillation in Local Image Editing

Abstract

Large instruction-based image editors can teach compact models, but local editing requires changing the target region while preserving its surroundings. Standard on-policy distillation (OPD) matches teacher and student at student-visited states, yet a global teacher target may propagate unintended changes outside the edit region. Uniform sampling across denoising states also ignores differences in their influence on the final edit. We introduce CSTAR-OPD, a Counterfactual Spatio-Temporal Alignment framework for Region-aware OPD. At the same student-visited state, we evaluate a frozen teacher under both the edit instruction and an identity instruction. The student matches the edit-conditioned velocity inside the target mask and the identity-conditioned velocity outside it, using separate region-normalized losses so that the background does not dominate the objective solely by area. The teacher and mask are used only during training; inference requires neither. Analyses of semantic direction and 4B-to-9B model switching motivate sampling one student state per example from states 0-13 of a 50-step schedule. On GIE-Bench, CSTAR-OPD reduces outside-mask M-MSE by 17.7% relative to full-range global OPD while maintaining the same 88.98% functional correctness. These results support combining early-state supervision with region-specific teacher targets to balance editing and preservation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.