CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
Quick Answer
CanvasAgent is a multimodal agent designed for complex image creation and editing, utilizing a new dataset called CanvasCraft, which includes 140K annotated trajectories.
Quick Take
It employs a hybrid reward system for training, enhancing its ability to manipulate visual states through multi-turn interactions. Experiments show that CanvasAgent effectively improves both image quality and workflow efficiency in multi-tool environments.
Key Points
- CanvasCraft dataset includes 140K fully annotated executable trajectories for training.
- CanvasAgent learns to orchestrate visual tools through multi-turn interactions.
- It combines outcome- and process-level signals for optimized training.
- Experiments demonstrate improved image quality and trajectory behavior.
- Existing multimodal agents lack large-scale supervision for complex image tasks.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing images, localizing objects, segmenting regions, editing selected content, compositing intermediate assets, reading text, and enhancing the final result. Such tasks shift multimodal agents from perception-augmented reasoning to manipulation-centered visual creation, where tools must actively transform visual states rather than merely inspect them. However, existing multimodal tool-use agents are mostly optimized for perception, search, or domain-specific editing, and lack large-scale supervision for executable image-creation trajectories. In this paper, we introduce CanvasCraft, a large-scale multimodal tool-use dataset for complex image creation and editing, and \textbf{CanvasAgent}, a tool-augmented multimodal agent that learns to orchestrate heterogeneous visual tools through multi-turn interaction. CanvasCraft contains 140K fully annotated executable trajectories and 10K
RL task specifications. CanvasAgent is first trained with SFT to learn executable reasoning-action trajectories, and is then optimized with GRPO using a hybrid reward that combines outcome- and process-level signals. During rollout, CanvasAgent inspects intermediate results, tracks visual assets, and adapts tool decisions to the evolving visual state. Experiments evaluate both final image quality and trajectory behavior, demonstrating the effectiveness of CanvasAgent and the proposed dataset for complex multi-tool image creation workflows.
| Comments: | 18pages, 5 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.05465 [cs.CV] |
| (or arXiv:2607.05465v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05465 arXiv-issued DOI via DataCite |
Submission history
From: HaiRui Zhu [view email]
[v1]
Mon, 6 Jul 2026 04:57:18 UTC (2,477 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.