VEvo-Harness: Causally-Guarded Self-Evolving Harness for Agentic Image Generation
Abstract
The central bottleneck in visual generation and editing has migrated from achieving pixel-level fidelity to satisfying densely constrained user specifications through increasingly agentic workflows. Prevailing agentic generation pipelines remain shallow generate-verify-rewrite loops treating VLM judgments as coarse visual feedback, either discarding repair experience or admitting ungated updates that poison skill library. With memorize and govern periods left open, they degrade into ungoverned iterators of repeated trial-and-error rather than continual self-evolving harness. Towards this issue, we present VEvo-Harness, the first causally-guarded continual self-evolving harness system for agentic visual generation. VEvo-Harness introduces four coupled innovations: (1) a visual traceback that densifies sparse judgments into root-cause attributions with a hierarchical verification tree; (2) a visual continual governance that admits only causally valid skill patches through counterfactual dual-arm image generation replay; (3) a lifecycle-managed harness state that hierarchically distills root-cause-anchored episodes across working, experience, and pattern tiers into a persistent skill library of routable and causally verified capabilities; (4) a heterogeneous sandbox policy that orchestrates a diverse action space of repair primitives, dynamically routing cost-aware interventions based on root causes to select the most effective strategy. VEvo-Harness surpasses state-of-the-art agentic methods and rivals proprietary generators (e.g., GPT-Image, Nano-Pro) while maintaining the continual harness evolution and generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.