MasterStudyBench: Can Agents Paint a Painting, or Only Replay Its Pixels?
Abstract
Genuine visual intelligence transcends passive pixel imitation: true mastery requires internalizing the generative hierarchy of visual creation. In classical art, an apprentice ascends through master studies, deconstructing a masterpiece's procedural logic rather than mechanically tracing its surface. Yet current visual code synthesis remains trapped in a pixel-only regime: grading solely on rendered appearance treats code as an uninspected black box, inviting reward hacking while remaining blind to painterly aesthetics and structural modularity alike. To bridge this chasm, we introduce MasterStudyBench, evaluating procedural artistry across 200 museum masterpieces inside an offline hardened sandbox. MasterStudyBench decouples two evaluation pillars: (1) an image-conditioned rubric tree of 17,323 criteria auditing fine-grained aesthetics across brushwork, lighting, and composition; and (2) an executed editability probe measuring re-rendered spatial deltas under counterfactual source mutations to test program modularity. Across frontier multimodal models, a profound disconnect emerges: adversarial stubs capturing nearly 67% pixel rewards collapse to 0.0 under MasterStudyBench, while rendered visual fidelity correlates only weakly with executed editability (Spearman's ), proving that visually convincing paintings routinely conceal uneditable spaghetti code. Mechanistically, architectural planning consistently yields a 2.6-4.7% editability gain by enforcing scene decomposition, while painterly aesthetics lags 11-16% behind basic visual integrity, exposing procedural shaders as the critical capability frontier. By separating painting from replaying, MasterStudyBench charts the path from superficial mimicry toward authentic procedural intelligence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.