MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion
Quick Answer
MGDT introduces a novel MLLM-Guided Diffusion Transformer for Multimodal Knowledge Graph Completion, enhancing entity inference by utilizing a Relation-Adaptive Mixture-of-Experts for semantic routing.
Quick Take
Experiments demonstrate that MGDT significantly outperforms existing methods across three benchmark datasets, addressing issues of noisy and inconsistent multimodal feature integration.
Key Points
- MGDT employs Relation-Adaptive Semantic Routing Mixture-of-Experts for effective cue selection.
- Utilizes a frozen Multimodal as a semantic anchor for alignment.
- Achieves superior performance on three benchmark datasets compared to strong baselines.
- Addresses challenges of noisy multimodal features in existing MKGC methods.
- Implements an align-then-diffuse paradigm for improved entity generation.
DeepSignal Analysis
What happened
MGDT is a new framework designed for Multimodal Knowledge Graph Completion (MKGC). It employs a Relation-Adaptive Mixture-of-Experts for selecting relevant multimodal features and uses a frozen Multimodal Large Language Model for semantic alignment. Experiments indicate that MGDT outperforms existing methods on three benchmark datasets.
Key evidence
- MGDT utilizes a Relation-Adaptive Semantic Routing Mixture-of-Experts module to enhance multimodal semantic transformation paths.
- The framework employs a frozen Multimodal Large Language Model as a semantic anchor to align multimodal representations.
- Experiments conducted on three benchmark datasets show that MGDT consistently outperforms strong baseline methods.
Why it matters
The introduction of MGDT addresses significant challenges in MKGC, particularly the issues of noisy and inconsistent multimodal feature integration. By improving the selection and alignment of multimodal cues, MGDT aims to enhance the accuracy of entity inference, which is crucial for applications in AI and knowledge representation.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues. Existing diffusion-based MKGC methods usually denoise directly on raw multimodal features. Such a design forces the denoiser to simultaneously perform relation-dependent cue selection, cross-modal semantic alignment, and structure-aware entity generation, which introduces noisy and semantically inconsistent conditions for diffusion and consequently leads to suboptimal completion performance. To address this limitation, we propose MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts (MGDT), a novel MKGC framework built on an align-then-diffuse paradigm. MGDT first employs a Relation-Adaptive Semantic Routing Mixture-of-Experts (RASR-MoE) module to select relation-relevant multimodal semantic transformation paths and suppress irrelevant modality interference. MGDT then uses a frozen Multimodal Large Language Model (MLLM) as a semantic anchor to align the routed multimodal representations into a unified latent space and reduce cross-modal semantic heterogeneity. Finally, a Knowledge Graph Diffusion Transformer (KGDT) performs graph-conditioned denoising generation in the aligned space to produce the missing entity representation. Experiments on three benchmark datasets show that MGDT consistently outperforms strong baselines.
| Comments: | 8pages, 6 figures |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.15592 [cs.AI] |
| (or arXiv:2607.15592v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15592 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xu Hou [view email]
[v1]
Fri, 17 Jul 2026 03:39:43 UTC (1,946 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.