acceptodds
Under review as a conference paper at ICLR 2027

ScalingEdit: Scaling Real-World 3D Object Editing with Self-Supervised Training

Abstract

3D object editing aims to rotate, scale, and translate a target object while preserving its identity and surrounding scene. Scaling up this capability is hindered by the scarcity of real-world editing pairs with geometric supervision. Existing multi-view videos offer diverse object observations and camera poses, but camera motion changes both the object viewpoint and the background, creating a mismatch with scene-preserving editing. To address these challenges, we introduce ScalingEdit, a self-supervised framework that turns these observations into training data for geometry-controlled object editing. Specifically, we formulate object editing as a foreground transformation within a shared scene context. We construct scene-aligned training pairs by compositing the source foreground onto the completed target-view background while retaining the unmodified target image as supervision. This construction aligns training with the scene-preserving editing task and enables learning from real multi-view observations without dedicated captures of object manipulation. A geometry-conditioning module encodes relative camera pose and image-plane displacement to provide explicit control over object rotation, scaling, and translation. Experiments on ObjectMover-A and our RotationBench demonstrate improvements over the evaluated baselines in appearance fidelity and geometric accuracy. Our data-scaling experiments further demonstrate the potential to improve geometric object editing by increasing real-world multi-view supervision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.