Detecting Hallucinations for Large Language Model-based Knowledge Graph Reasoning
Quick Answer
The paper introduces LUCID, a novel hallucination detection method for LLM-based knowledge graph reasoning, which integrates LLM attention scores, KG semantics, and structural information.
Quick Take
Experiments demonstrate that LUCID outperforms 15 baselines across nine datasets, addressing critical misinformation issues in existing frameworks.
Key Points
- LUCID is the first method specifically designed for hallucination detection in KG reasoning.
- It utilizes attention scores, semantic similarities, and KG structure via a graph neural network.
- The method was evaluated on nine benchmark datasets, achieving state-of-the-art results.
- LUCID addresses the limitations of existing detection methods that ignore KG structural information.
- This advancement is crucial for improving the reliability of -based decision support systems.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Knowledge graph (KG) reasoning infers new knowledge from existing facts and is widely applied in question answering, recommendation, and decision support. With the rapid development of large language models (LLMs), LLM-based KG reasoning frameworks have become increasingly popular by leveraging retrieved KG information. However, hallucinations in LLMs remain a critical issue. Even when relevant KG knowledge is incorporated, models may still generate incorrect outputs, leading to misinformation and unreliable decisions. Existing hallucination detection methods either focus on LLM internal states or verify consistency with retrieved contexts, but both overlook the structural information in KGs, resulting in suboptimal performance. To address this gap, we propose LUCID, the first halLUcination deteCtIon method for LLM-based knowleDge graph reasoning frameworks. LUCID jointly leverages LLM attention scores, KG semantics, and structural information. Specifically, it extracts node and edge features from attention scores and semantic similarities, and integrates them with KG structure using a graph neural network. We also construct manually annotated benchmark datasets for evaluation. Experiments on nine datasets show that LUCID achieves state of the art performance compared to 15 baselines.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.19351 [cs.CL] |
| (or arXiv:2606.19351v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.19351 arXiv-issued DOI via DataCite |
Submission history
From: Xinyan Zhu [view email]
[v1]
Mon, 27 Apr 2026 12:20:28 UTC (773 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.