acceptodds
Under review as a conference paper at ICLR 2027

3D Code Factory: Scaling Verifiable 3D Tasks for Coding Agents

Abstract

Recent rapid progress in coding agents has increased the demand for scalable task generation to advance agent capabilities, which requires diverse, executable tasks equipped with reliable automated verifiers. For procedural 3D modeling, fixed-reference comparisons can penalize semantically valid alternative implementations, while rendered images alone cannot fully verify spatial relationships and precise measurements. To address this problem, we introduce 3D Code Factory, an automated framework that synthesizes executable 3D coding tasks paired with multimodal evaluation rubrics from seed scenarios, decomposing verification into scene-state checks for spatial relations and rendered-image checks for visual appearance. Building on this framework, we construct 3D Live Bench to evaluate coding agents on 36 core tasks in Three.js and Blender under different scaffolds, assessing deliverables across three metrics: execution success, rubric compliance, and visual quality. Extensive evaluations reveal substantial room for improvement in complex 3D modeling: despite high program execution rates, agents struggle to satisfy coupled task requirements, with peak visual quality and rubric compliance achieved by different model–scaffold configurations. During iterative refinement, localized repairs often introduce regressions, invalidating or neglecting previously satisfied criteria. Furthermore, scaffold design fundamentally alters how base-model capabilities translate into task completion, while prolonged edit–render loops substantially inflate token expenditure with diminishing returns.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.