An Explainable Header-Centric Framework for Large-Scale Semantic Table Interpretation and Data Quality Assessment
Quick Answer
This paper introduces an explainable, header-centric framework for metadata-only Column Type Annotation and Data Quality Assessment, effectively mapping headers to 39 FinalFormat types.
Quick Take
Evaluated on diverse benchmarks with around 120,000 header columns, it demonstrates broad applicability in noisy real-world metadata, although initial evaluations show modest performance due to granularity and ontology selection issues.
Key Points
- Framework maps headers to 39 interpretable FinalFormat types using curated lexical resources.
- Detects data quality issues like missing data, duplicates, and wrong data types.
- Evaluated on benchmarks including UCI, Kaggle, and SemTab 2024 with 120,000 headers.
- Results indicate broad practical coverage in noisy real-world metadata.
- Provides a reusable workflow for metadata-driven semantic annotation and quality monitoring.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Knowledge Graph (KG) quality depends not only on downstream graph validation, but also on the quality of tabular metadata used before integration. In metadata-only Semantic Table Interpretation (STI), where cell values are unavailable, noisy, or unsuitable, column headers become a critical source of semantic evidence for traceable KG preparation.
We present an explainable, header-centric framework for metadata-only Column Type Annotation (CTA) and Data Quality Assessment (DQA). The framework maps headers to 39 interpretable FinalFormat types using curated lexical resources and preserves token-level traceability through SourceKeywords. Each assigned type activates validation rules based on a taxonomy of Data Quality Issues (DQIs), producing detections such as missing data, duplicates, domain violations, wrong data type, and temporal mismatch. These detections are aggregated into HeadersIQ, a lightweight, unweighted data source-level quality metric.
The framework was evaluated across heterogeneous benchmarks, including UCI, Prague, Kaggle, VizNet/Sato, SOTAB, T2Dv2, and the SemTab 2024 Metadata-to-KG track, comprising around 120,000 header columns. The results show broad practical coverage across noisy real-world metadata, while a parallel KG-mapping pathway supports alignment to DBpedia and this http URL. On the SemTab 2024 Metadata-to-KG track, the official GT-strict evaluation was modest. However, a blinded diagnostic audit indicates that many mismatches reflect benchmark granularity, aliasing, and ontology-selection effects rather than wholly implausible header-centric predictions. We report this audit as diagnostic evidence on disagreement patterns, not as revised benchmark performance. Overall, the paper presents a reusable workflow for metadata-driven semantic annotation, data source-level quality monitoring, and KG-oriented benchmark diagnosis.
| Comments: | 18 pages, 4 figures, Workshop on Quality of Knowledge Graphs at ESWC 2026, May 11, 2026, Dubrovnik, Croatia |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.10541 [cs.AI] |
| (or arXiv:2610.10541v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10541 arXiv-issued DOI via DataCite |
Submission history
From: Marcelo Valentim Silva [view email]
[v1]
Tue, 12 May 2026 05:52:55 UTC (218 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.