XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
Quick Answer
XL-DocBench introduces a human-verified benchmark for extra-long document understanding, featuring 1,519 questions across six domains and contexts up to 2,303 pages.
Quick Take
It emphasizes multi-page evidence and structured reasoning, revealing that current systems struggle with long contexts and complex document interactions.
Key Points
- XL-DocBench includes 1,519 questions from six professional domains.
- 72.6% of questions require evidence from multiple pages.
- 36.6% of questions involve tables, charts, or figures.
- The benchmark was verified by 194 human experts.
- Current systems struggle with long contexts and structured reasoning.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Hongchen Wei, Yuanzhe Wang, Bei Liu, Yifan Yang, Qi Dai, Ruichun Ma, Kai Qiu, Yunsheng Li, Dongdong Chen, Chong Luo, Zhenzhong Chen, Baining Guo
Abstract:Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages. Some questions also require comparing related reports. Reliable long-document understanding is therefore a prerequisite for using LLMs in compliance, clinical, financial, and engineering workflows, where decisions must be traceable to specific evidence pages and the cost of an unsupported answer is high -- yet most existing benchmarks still measure short-context or single-page QA. We introduce XL-DocBench, a fully human-verified benchmark for extra-long document understanding, with 1,519 retained questions from six professional domains and contexts up to 2,303 pages. XL-DocBench goes beyond page-level lookup. 1,103 examples (72.6\%) use multiple evidence pages. The final set also includes 556 questions (36.6\%) that use tables, charts, or figures, and 165 questions (10.9\%) that require evidence from multiple documents. Each question has one of twelve reasoning labels, expert-annotated evidence pages, a typed verification rule, and an answer format, including 218 None-answer cases. We build the benchmark with a tree-guided synthesis pipeline followed by artifact filters and full verification by 194 human experts. By coupling extra-long professional contexts with page-level evidence and typed rules, XL-DocBench fills a gap left by prior single-page, short multi-page, or text-only long-context benchmarks, and lets future work attribute system failures to retrieval, evidence use, or rule following rather than to a single leaderboard score. The results show that current systems still struggle with long contexts, multi-page evidence, and structured reasoning over professional documents.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2608.00036 [cs.CL] |
| (or arXiv:2608.00036v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2608.00036 arXiv-issued DOI via DataCite |
Submission history
From: Hongchen Wei [view email]
[v1]
Tue, 21 Jul 2026 09:38:24 UTC (13,313 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.