An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents
Quick Answer
This paper introduces an ontology-guided extraction layer for knowledge graph construction, utilizing a Qwen3.5-9B model to enhance entity and relationship extraction from heterogeneous documents.
Quick Take
The system achieves a 94% reduction in catalog overhead and improves search recall from 70% to 95% without false merges, addressing various quality defects in the data.
Key Points
- Utilizes Qwen3.5-9B model for entity and relationship extraction.
- Reduces catalog overhead by 94% using live ontology retrieval.
- Improves search recall from 70% to 95% with no false merges.
- Includes five-stage refinement pipeline for data quality.
- Addresses silent quality defects, enhancing overall extraction accuracy.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct individuals who share a name risk silent conflation. This paper presents the design, implementation, and empirical refinement of a production extraction layer that converts a live document stream into a validated knowledge graph aligned to a formal ontology. The system consumes document metadata from Kafka, routes PDF, spreadsheet, Office, and image content through handlers built for each format, and extracts entities and relationships in two passes using a locally hosted Qwen3.5-9B model tuned on the ontology. Its distinguishing component is ontology-guided extraction: the relevant slice of a curated ontology is retrieved live from a graph database by embedding similarity and injected into the extraction prompt, reducing catalog overhead by about 94 percent relative to static domain slices. Extracted results then pass through a refinement pipeline of five stages: deterministic cleaning, merging across chunks, a second pass for relationships, six deduplication algorithms that require no model inference, and an embedding resolution subsystem whose conflict guard no similarity score can override. Evaluation on intelligence corpora improved search recall from roughly 70 to 95 percent with no false merges, and corrected seven classes of silent quality defect, ranging from a bug that truncated source text by a single character to the systematic duplication of entities that carried title prefixes.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.28662 [cs.AI] |
| (or arXiv:2607.28662v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.28662 arXiv-issued DOI via DataCite |
Submission history
From: Vaibhav Dangaich [view email]
[v1]
Wed, 22 Jul 2026 09:45:26 UTC (68 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.