Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies
Quick Answer
LoFa introduces a benchmark for assessing LLM robustness against logical fallacies, revealing varying vulnerability profiles among models.
Quick Take
The proposed metric, LFR@k, quantifies resistance to fallacious arguments, highlighting the need for improved resilience in .
Key Points
- LoFa benchmarks LLM resilience against logical fallacies through a pipeline.
- The framework includes a multi-round debate to test model robustness under persuasion.
- LFR@k metric quantifies logical fallacy resistance, addressing knowledge limitations.
- Experiments show LLMs have varying robustness across different types of fallacies.
- Distinct vulnerability profiles among models indicate specific areas for improvement.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies remains underexplored. Prior work has primarily examined whether LLMs can identify or classify fallacies, leaving their robustness against fallacious persuasion insufficiently studied. To address this gap, we introduce LoFa (Logical Fallacy), a comprehensive benchmark for evaluating LLM robustness against fallacies. LoFa is constructed through a multi-agent pipeline that pairs factual questions with fallacious arguments, and is accompanied by a multi-round debate framework for assessing model resilience under sustained adversarial persuasion. To disentangle fallacy robustness from a model's inherent knowledge limitations, we further propose Logical Fallacy Resistance at k (LFR@k), a metric that quantifies resistance to fallacious attacks. Experiments show that LLMs exhibit varying levels of robustness across different fallacy types, revealing distinct vulnerability profiles among models.
| Comments: | Accepted to ACL 2026 Main. 33 pages (9 pages main text) |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2606.31039 [cs.CL] |
| (or arXiv:2606.31039v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.31039 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xudong Shen [view email]
[v1]
Tue, 30 Jun 2026 02:17:45 UTC (2,038 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.