DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification
Quick Answer
This paper shows that The DS@GT ARC team developed a system for the CLEF 2026 CheckThat!
Quick Take
Task 2, utilizing for multilingual numerical claim verification. Their LLM-based approach outperformed a TF-IDF reward model in most metrics, particularly Recall@5, while AraBERT excelled over a multilingual model for Arabic claims.
Key Points
- LLM-based verifier uses LoRA for independent trace scoring and Best-of-N selection.
- Adaptive sub-claim decomposition did not enhance performance, introducing noise instead.
- AraBERT outperformed multilingual models on most metrics for Arabic claims.
- The LLM approach excelled in Recall@5, while the reward model was better on Conflicting class.
DeepSignal Analysis
What happened
The DS@GT ARC team presented a system for verifying numerical claims in English and Arabic at the CLEF 2026 CheckThat! Task 2. Their LLM-based approach outperformed a TF-IDF reward model in most metrics, particularly in Recall@5, while AraBERT showed superior performance for Arabic claims.
Key evidence
- The system focuses on ranking reasoning traces generated by large language models (LLMs) and predicting final verdicts for numerical claims in English and Arabic.
- The LLM-based verifier was fine-tuned using LoRA to independently score reasoning traces, while a lightweight TF-IDF model used handcrafted features for scoring.
- AraBERT outperformed a general multilingual model across most metrics for Arabic claims, indicating the effectiveness of language-specific models.
Why it matters
This research highlights the challenges of automated verification of numerical claims, which requires both language understanding and quantitative reasoning. The findings suggest that LLMs can significantly enhance performance in multilingual contexts, particularly for complex tasks like claim verification. The results also indicate that language-specific models may provide better accuracy than general multilingual models, which could influence future developments in natural language processing and AI applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Automated verification of numerical claims is a challenging problem, as it requires both language understanding and quantitative reasoning. This paper describes our system for CLEF 2026 CheckThat! Task 2, which focuses on ranking reasoning traces generated by large language models (LLMs) and predicting a final verdict for numerical claims in English and Arabic. We explore two approaches. The first approach fine-tunes an LLM-based verifier using LoRA to score each reasoning trace independently as a binary classification problem, and selects the final verdict using Best-of-N selection. We further experiment with adaptive sub-claim decomposition to break complex claims into simpler parts before verification. The second approach uses a lightweight TF-IDF reward model with handcrafted numeric and temporal overlap features to score traces, and aggregates scores by verdict group to determine the final prediction. For Arabic, we compare a general multilingual model against AraBERT, a language-specific model pretrained on Arabic text. Our results show that the LLM-based approach outperforms the lightweight reward model on most metrics, particularly Recall@5, while the reward-based approach shows stronger performance on the Conflicting class. Sub-claim decomposition did not improve performance, suggesting that claim splitting introduces noise rather than aiding reasoning. For Arabic, AraBERT outperforms the multilingual baseline across most metrics.
| Comments: | 10 pages, 2 figures. Accepted at CLEF 2026 CheckThat!. To appear in CEUR Workshop Proceedings |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.25069 [cs.CL] |
| (or arXiv:2607.25069v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.25069 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sagnik Sinha [view email]
[v1]
Mon, 27 Jul 2026 20:59:14 UTC (159 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.