ARGUS: An Evolving, Auditable Evaluator for 3D Scene Editing Agents via Evidence-Driven Skill Induction
Abstract
A 3D scene editing agent can produce a plausible scene without completing the requested edit, so evaluators that accurately check the executed result are increasingly important. Image-based judgment often fails when the target is occluded or the change is small, and an aggregate score or binary verdict discards which conditions were checked, whether a checker existed, and whether the evidence sufficed, which is the information repair needs. We propose ARGUS, an experience-driven framework whose primary objective is accurate evaluation and whose downstream objective is scene improvement. It compiles a task specification, binds each requirement to compatible checkers, and combines state, visual, physical, and preservation evidence; confirmed violations go to repair, missing evidence to additional observation, and missing capability to skill learning. A separate pipeline has an LLM propose executable evaluation and repair programs from reviewed evidence. Admission tests each candidate on labeled counterexamples and on a label-free invariance test that replays it under meaning-preserving transformations such as unit conversion and world-frame rotation, which in a controlled test rejects faulty measurements that single-axis labels admit; only programs that pass both become reusable skills. On 48 edits where every evaluator sees only the instruction and the scenes, ARGUS has the lowest selective risk (error rate among decided conditions), 7.0%, and the highest scene AUROC, 81.9, and it holds conditions it cannot settle with the missing evidence named. Registered skills, together with a fixed relation frame in the binder, raise coverage (the fraction of conditions decided) from 44.2% to 60.0% at the same error rate; the skills alone decide 11 held conditions, 10 correctly. Accurate diagnosis carries over to repair, raising the holistic success rate from 56.8% to 65.1%, and admitted skills are reused without LLM calls and revised when new counterexamples appear.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.