Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science
Quick Answer
The ArcticQA dataset introduces 194 Arctic science questions to evaluate LLM abstention, revealing that abstention rates vary significantly across models like Gemini and ChatGPT.
Quick Take
The study shows that replacing correct answers with distractors increases abstention by an average of 5.05 percentage points, emphasizing the need for a joint evaluation of abstention frequency and responsiveness.
Key Points
- ArcticQA consists of 194 questions from primary Arctic research.
- Eight models, including Gemini, Claude, and ChatGPT, were evaluated.
- Abstention rates for answer-present conditions ranged from 0.0% to 63.0%.
- Replacing correct answers with distractors increased abstention by 5.05 percentage points.
- The dataset and benchmark are publicly available for further research.
DeepSignal Analysis
What happened
The ArcticQA dataset was introduced to assess large language models' (LLMs) abstention rates in Arctic science questions. It includes 194 questions and evaluates eight models, revealing abstention rates ranging from 0.0% to 63.0%. The study found that replacing correct answers with distractors increased abstention by an average of 5.05 percentage points.
Key evidence
- The ArcticQA dataset consists of 194 questions derived from primary Arctic research, with automated checks for answer support.
- Eight models from the Gemini, Claude, and ChatGPT families were evaluated, resulting in 9,312 recorded responses across three trials per condition.
- The study found that answer-present abstention rates varied significantly, with a range from 0.0% to 63.0%.
Why it matters
Understanding LLM abstention is crucial for evaluating their reliability in scientific contexts. The significant variation in abstention rates across models indicates differing capabilities in handling scientific questions. The findings underscore the importance of assessing both abstention frequency and responsiveness to improve model performance in real-world applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at this https URL.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.09446 [cs.CL] |
| (or arXiv:2610.09446v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09446 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yunhe Feng [view email]
[v1]
Wed, 7 Oct 2026 04:58:28 UTC (87 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.