Long Live Fine-Tuning: Task-Specific Transformers Outperform Zero-Shot LLMs for Misinformation Response Classification on Reddit
Quick Answer
This paper shows that Fine-tuned RoBERTa outperforms zero-shot models like Claude Haiku 4.5 in misinformation classification on Reddit, achieving a macro-F1 of 0.62 versus 0.50.
Quick Take
This highlights that task-specific tuning is crucial for detecting belief, a category often missed by larger models. Despite the rise of large , fine-tuning remains the more effective approach for nuanced tasks.
Key Points
- Fine-tuned RoBERTa achieves 0.62 macro-F1, outperforming Claude Haiku 4.5's 0.50.
- Llama-3-8B's performance matches Llama-3-70B, indicating scaling doesn't guarantee better results.
- Zero-shot models struggle with belief detection, a critical aspect in misinformation classification.
- Task-specific fine-tuning is more cost-effective and reliable for nuanced classification tasks.
- Label schema and topic significantly influence zero-shot model performance.
Paper Resources
Article Content
From source RSS / original summaryarXiv:2606. 04274v1 Announce Type: new Abstract: As (LLMs) become default tools for online information verification, an implicit assumption follows them: that scale and general capability are sufficient for nuanced classification of misinformation discourse. We test this assumption directly on 900 Reddit comments spanning three PolitiFact-verified misinformation claims (environment, health, immigration), labelled as belief (propagates the claim), fact-check (corrects it), or other.
We compare nine models across three paradigms -- BART-MNLI, three Llama variants, three commercial frontier LLMs (Claude Haiku 4. 5, Gemini Flash Lite 2. 5, Claude Sonnet 4. 6), and fine-tuned DistilBERT and RoBERTa -- under universal and topic-specific label schemas. The assumption does not hold. Fine-tuned RoBERTa reaches 0. 62 macro-$F_1$ against a best zero-shot result of 0. 50 (Claude Haiku 4.
5), at a fraction of the per-query cost; the supervised advantage is concentrated on the belief class, the implicit, affective category every zero-shot model under-detects. Scaling does not help: Llama-3-8B matches Llama-3-70B, and Claude Sonnet 4. 6 underperforms the smaller Haiku under generic labels, collapsing belief detection to 0. 17 and refusing outright on a subset of comments flagged as sensitive. This is a safety-alignment artefact, not a capacity limit.
Label schema and topic jointly shape zero-shot performance, with the same model varying by more than 0. 13 macro-$F_1$ across topics under matched labels. In a verification context, where missing belief is the costlier error, task-specific fine-tuning remains the more reliable choice despite the proliferation of large generative models.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.