acceptodds
Under review as a conference paper at ICLR 2027

CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration

Abstract

Complex image creation and editing often exceed a single generation or editing model: a request may require synthesizing, localizing, segmenting, editing, compositing, reading text, and enhancing within one workflow. Such requests are complex in three respects: they involve long-horizon, state-dependent execution, heterogeneous tools whose outputs feed subsequent operations, and multiple visual assets, including multiple input images. Existing multimodal tool-use agents mostly target perception, search, or domain-specific editing, or plan over a single image without training, and large-scale supervision for executable, multi-asset creation trajectories remains scarce. We introduce CanvasCraft, a large-scale multimodal tool-use dataset, and CanvasAgent, an agent that learns to orchestrate heterogeneous visual tools while tracking intermediate visual assets. CanvasCraft contains fully annotated executable trajectories and RL task specifications stratified by reasoning difficulty, trajectory length, and tool diversity. CanvasAgent is trained with SFT and then GRPO, using a hybrid reward that combines outcome- and process-level signals. On CanvasCraft-Test (250 tasks), CanvasAgent raises the overall reward of its Qwen3-VL-8B base from 0.426 to 0.821, improving both final-image alignment and trajectory quality, and human evaluation shows the same trend.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.