Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers
Quick Answer
This paper shows that A new study introduces Symbolic Augmentation to enhance the robustness of neural fact-checkers against canonical-equivalence errors, achieving a macro-F1 score of 0.902 on a 1500-item benchmark.
Quick Take
This method significantly improves accuracy from 36.5% to 98.2% for physically equivalent quantities while maintaining performance on external benchmarks like SciFact-Open. The findings highlight the importance of training-time augmentation in integrating symbolic and learned components.
Key Points
- Introduced a five-class typed-quantity error taxonomy for fact-checking.
- ModernBERT fine-tuned on the benchmark achieved a macro-F1 of 0.899.
- Symbolic Augmentation improved canonical-equivalence accuracy from 36.5% to 98.2%.
- Augmented encoder matched a closed-frontier without inference cost.
- Negative results indicated no benefit from symbolic features as auxiliary inputs.
DeepSignal Analysis
What happened
A study introduces Symbolic Augmentation to improve neural fact-checkers' accuracy against canonical-equivalence errors. This method increased accuracy for physically equivalent quantities from 36.5% to 98.2% and achieved a macro-F1 score of 0.902 on a 1500-item benchmark. The findings suggest that training-time augmentation is crucial for integrating symbolic and learned components.
Key evidence
- The study developed a five-class typed-quantity error taxonomy and a 1500-item benchmark, achieving a Krippendorff's alpha of 0.882 for labeling accuracy.
- A ModernBERT encoder fine-tuned on the benchmark achieved a macro-F1 score of 0.899, outperforming existing neural fact-checkers.
- Symbolic Augmentation improved the encoder's robustness to canonical-equivalence errors from 36.5% to 98.2% while slightly enhancing in-distribution accuracy.
Why it matters
The research addresses a significant blind spot in neural fact-checkers related to canonical-equivalence errors, which can lead to incorrect scientific claims. By enhancing accuracy through Symbolic Augmentation, the study provides a potential pathway for improving the reliability of automated fact-checking systems. This advancement is particularly relevant in the context of scientific communication, where precision is critical.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models hallucinate numbers and units when summarizing scientific text, a failure mode that can silently invert a scientific claim. We recast the detection of such errors as typed verification: we introduce a five-class typed-quantity error taxonomy and a 1500-item benchmark, rewritten from PMC and arXiv sources and labeled by two independent LLM annotators with adjudication (Krippendorff's alpha = 0.882). A ModernBERT encoder fine-tuned on this benchmark reaches macro-F1 = 0.899, far above any off-the-shelf neural fact-checker, yet four probes expose a sharp structural blind spot: on canonical-equivalent rewrites of physically equivalent quantities (e.g., 95°C and 368.15 K) its accuracy collapses to 36.5%. We propose Symbolic Augmentation, a training-time framework that runs the modules of a symbolic verifier in reverse to generate label-preserving augmented training data. The augmentation lifts canonical-equivalence robustness to 98.2% while slightly improving in-distribution accuracy (macro-F1: 0.899 to 0.902); the augmented encoder matches a closed-frontier LLM at no inference cost and transfers to an external benchmark (SciFact-Open binary macro-F1: 0.791 to 0.828). Two negative results sharpen the claim: symbolic features as auxiliary encoder inputs add nothing, and symbolic silver labels scale negatively under teacher noise. Together these results identify training-time augmentation as the right integration point between symbolic and learned components.
| Comments: | 18 pages, 3 figures, 6 tables |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.16212 [cs.AI] |
| (or arXiv:2607.16212v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16212 arXiv-issued DOI via DataCite |
Submission history
From: Genpei Zhang [view email]
[v1]
Tue, 19 May 2026 15:58:00 UTC (315 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.