acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking and Evolving the Capability Boundary in Agentic Visual Generation

Abstract

Visual generators render complex scenes convincingly, yet they confidently fabricate the characters, symbols, and facts they never learned. This world-knowledge bottleneck is structural: generators learn from fixed corpora, whereas user requests are unbounded, evolving, and long-tailed. To measure and study it, we construct the first large-scale dataset SearchGen-20K, SearchGen-Corpus-1M and SearchGen-Bench featuring 20,937 production-level generation tasks paired with truth-grounded verifiers that check the requested content rather than image quality alone. Comparing open-weight generators with agent-powered frontier APIs, our analysis reveals a substantial gap of up to 35pp in handling production-level requests. Agentic visual generation, in which an agent retrieves evidence for the generator, is the natural remedy. Yet we identify two challenges: naive search can override generator capabilities due to search noise, and agent become stale when generator evolves its capability. In response, we showcase a minimal baseline approach that introduces a noise-resistant agentic protocol and co-adapts the generator and the agent. Data, code, and models are released for future research.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.