acceptodds
Under review as a conference paper at ICLR 2027

UniVEdit: Unified Visual Editing across Images, Videos, and 3D Scenes

Abstract

Recent advances in unified visual editing have begun to consolidate image and video editing into a single model, yet 3D scene editing remains largely isolated. A major challenge for 3D editing is the scarcity of high-quality paired training data, which limits the learning of generalizable editing capabilities. Meanwhile, image, video, and 3D editing share common operations and editing semantics, making knowledge transfer across the three modalities possible. Motivated by this observation, we present UniVEdit, a unified framework for image, video, and 3D scene editing within a single model. UniVEdit unifies the three modalities through a shared sequence representation and progressive multimodal training, allowing abundant image and video editing knowledge to benefit data-scarce 3D editing. However, shared editing semantics alone do not guarantee geometrically consistent outputs. We therefore introduce geometry-aware post-training, using a pretrained geometry reconstruction model to supervise depth and camera consistency. We further investigate how geometric supervision should be applied across modalities, and find that applying it only to 3D samples yields the best overall results. Under this setting, although image and video editing receive no direct geometric supervision, they also show gains, suggesting that geometric regularization learned from 3D data can transfer through the shared backbone. Further analysis shows that image and video training mainly transfers editing semantics to 3D, while limited 3D training primarily improves multi-view consistency; geometry-aware post-training further improves overall 3D editing performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.