CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
Abstract
Complex image creation and editing often exceed a single generation or editing model: a request may require synthesizing, localizing, segmenting, editing, compositing, reading text, and enhancing within one workflow. Such requests are complex in three respects: they involve long-horizon, state-dependent execution, heterogeneous tools whose outputs feed subsequent operations, and multiple visual assets, including multiple input images. Existing multimodal tool-use agents mostly target perception, search, or domain-specific editing, or plan over a single image without training, and large-scale supervision for executable, multi-asset creation trajectories remains scarce. We introduce CanvasCraft, a large-scale multimodal tool-use dataset, and CanvasAgent, an agent that learns to orchestrate heterogeneous visual tools while tracking intermediate visual assets. CanvasCraft contains fully annotated executable trajectories and RL task specifications stratified by reasoning difficulty, trajectory length, and tool diversity. CanvasAgent is trained with SFT and then GRPO, using a hybrid reward that combines outcome- and process-level signals. On CanvasCraft-Test (250 tasks), CanvasAgent raises the overall reward of its Qwen3-VL-8B base from 0.426 to 0.821, improving both final-image alignment and trajectory quality, and human evaluation shows the same trend.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.