acceptodds
Under review as a conference paper at ICLR 2027

3DE-Agent: Toward Universal Instruction-Driven Agentic 3D Editing

Abstract

Due to the lack of generalizable 3D editing protocols and paired cross-task 3D assets, existing instruction-driven 3D editing methods, based on either explicit representations or implicit models, are typically limited to simple and specific editing tasks, making it difficult to support diverse and complex editing intents through a universal instruction interface. In this paper, we propose 3DE-Agent, a universal agentic framework to perform various and composite 3D editing tasks. To achieve this goal, we first encapsulate off-the-shelf multi-modality models and 3D operations as optional tools, then build a 3D editing harness that enables VLM to fully exploit inherent reasoning and tool-use capabilities for understanding and performing versatile and complex editing instructions. In addition, we observe that agentic operation sometimes produces undesired regions, such as holes after object removal or blurs after deformation. To mitigate these issues, we curate pseudo paired 3D multi-view data to re-purpose Stable Virtual Camera (SEVA) as a general-purpose 3D refiner via on-policy self-distillation, improving both visual quality and cross-view consistency of editing results. Extensive experiments demonstrate the effectiveness and generalization of 3DE-Agent in tackling diverse and complex instructions. Codes and models will be publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.