Knowledge-Centric Agents for Workflow Generation
Quick Answer
The proposed knowledge-centric framework enhances workflow generation in visual systems like ComfyUI by modeling knowledge structures and dynamics.
Quick Take
It achieves superior results in node diversity, structural coherence, and execution success rates compared to existing approaches, establishing a new standard for agentic workflow generation.
Key Points
- Introduces a knowledge-centric framework for workflow generation in visual systems.
- Implements knowledge inversion to create hierarchical representations from real-world workflows.
- Utilizes supervised fine-tuning for knowledge injection, improving reasoning from tasks to strategies.
- Achieves higher execution success rates and richer node diversity than existing systems.
- Sets a new foundation for knowledge-driven workflow generation in AI applications.
DeepSignal Analysis
What happened
A new knowledge-centric framework for workflow generation in visual systems like ComfyUI has been proposed. This framework aims to improve upon existing large language model (LLM) approaches by incorporating knowledge structures and reasoning dynamics, leading to better outcomes in workflow generation.
Key evidence
- The framework addresses limitations in existing LLM approaches, which often struggle with structural brittleness in workflow generation tasks.
- Knowledge inversion is employed to create hierarchical representations from real-world workflows, enhancing the model's understanding of task structures.
- Experiments indicate that the proposed method achieves higher execution success rates and greater node diversity compared to current systems.
Why it matters
This development is significant as it establishes a new standard for agentic workflow generation, potentially transforming how visual systems operate. By focusing on knowledge modeling, the framework could lead to more effective and reliable workflows, which is crucial for applications requiring complex reasoning.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over modular compositions. Existing large language model (LLM) approaches often treat this as a direct text-to-JSON generation task, struggling with structural brittleness and lacking the experiential knowledge required for effective design. We argue that successful workflow generation requires modeling knowledge itself, including its structure, hierarchy, and reasoning dynamics. To this end, we propose a knowledge-centric framework that learns to invert, inject, and infer with knowledge across multiple abstraction levels. We first perform knowledge inversion to distill hierarchical representations, ranging from full pseudo-codes and skeletons to high-level strategies, from large collections of real-world workflows. We then conduct knowledge injection through supervised fine-tuning, teaching the model to reason from task descriptions to strategies and from strategies to executable structures. During inference, the model performs reversible reasoning to synthesize executable workflows, augmented by self-refinement for structural coherence. Extensive experiments demonstrate that our method produces workflows with richer node diversity, more coherent structures, and higher execution success rates than existing systems, establishing a new foundation for knowledge-driven, agentic workflow generation.
| Comments: | Accepted to ECCV 2026 |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.15845 [cs.AI] |
| (or arXiv:2607.15845v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15845 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Lei Sun [view email]
[v1]
Fri, 17 Jul 2026 11:03:01 UTC (11,804 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.