acceptodds
Under review as a conference paper at ICLR 2027

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

Abstract

3D editing methods are typically evaluated one edit at a time, although creating an asset requires a sequence of revisions. A method must follow each new instruction while preserving earlier edits and untouched regions, yet this ability has not been systematically measured. We introduce EditHero, to our knowledge the first benchmark for long-horizon 3D editing from image and natural-language instructions covering both geometry and texture. A deterministic assembly engine produces the target after every edit, and manual review verifies each sequence. Evaluating four 3D editing methods and six large language models (LLMs) reveals that conventional methods often miss the requested changes and alter regions meant to be preserved. Most LLMs that edit through code follow instructions more accurately, and all of them preserve unchanged regions well, but their time and token costs limit interactive use. We release the synthesis engine and edit sequences to support research on reliable iterative 3D editing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.