acceptodds
Under review as a conference paper at ICLR 2027

Point-Conditioned Drag Editing for In-Context Diffusion Transformers

Abstract

Point-based drag editing enables intuitive spatial control by allowing users to specify sparse source-target point pairs. However, converting such sparse geometric directives into coherent image edits remains challenging, especially when the edited content must move accurately while preserving identity, background, and local visual details. Existing drag-editing methods often rely on iterative point tracking, hand-crafted motion propagation, region-level assumptions, or mask-guided inpainting, which can introduce additional user burden and limit generalization to complex motion. We propose PiDiT, a point-conditioned drag-editing framework built on an in-context diffusion transformer. PiDiT renders sparse drag pairs into spatial point maps, encodes them with a lightweight point encoder, and injects the resulting token-aligned point features directly into the transformer's hidden states. This design treats point pairs as internal spatial conditions for the target generation branch while preserving the source image as in-context visual evidence. The framework is trained with a flow-matching objective and requires only the source image and sparse point pairs at inference time, without user-provided masks or test-time optimization. Experiments on DragBench and ReD-Bench show that PiDiT improves point reachability while maintaining competitive visual fidelity and identity preservation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.