COMPASS: COnditioned Manipulation via Precise Aligned Spatial Specification
Abstract
Image editing models are essential for content creation, and instruction-based editors make semantic modifications accessible through natural language. Precise spatial control (i.e., relocating, rotating, or rescaling an object) remains highly unreliable and challenging. We present COMPASS, which conditions a pretrained image editor on rendered 3D bounding-box proxies: we rasterize a proxy box at the source pose and another at the target pose, and paint each face with a fixed color shared across boxes. This allows the spatial control to be represented as local, spatially aligned differences in the pixel domain, which a diffusion transformer can follow without learning cross-position geometry from scratch. COMPASS also supports multi-object edits in a single pass. To demonstrate our approach, we use Qwen-Image-Edit as the base backbone and fine-tune it on a new dataset of 37,000 source/target image pairs, of which 30,000 are used for training and 7,000 are held out as a benchmark with single-object, multi-object, and overlapping-box splits; we further transfer the model zero-shot to the WildDet3D real-image dataset. We compare COMPASS with state-of-the-art editors on this benchmark. COMPASS outperforms existing techniques: (1) on the new dataset, COMPASS achieves a translation error of 0.27 m and a target-mask IoU of 0.765, where competing methods range from 0.31 to 1.19 m and from 0.262 to 0.633; and (2) on the real-image dataset, COMPASS reduces translation error to 4.15 m, nearly half of the best competing method at 7.93 m and under 40% of the worst at 11.83 m.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.