XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
Quick Answer
XL-DocBench introduces a human-verified benchmark for extra-long document understanding, featuring 1,519 questions across six domains and contexts up to 2,303 pages.
Quick Take
It emphasizes multi-page evidence and structured reasoning, revealing that current systems struggle with long contexts and complex document interactions.
Key Points
- XL-DocBench includes 1,519 questions from six professional domains.
- 72.6% of questions require evidence from multiple pages.
- 36.6% of questions involve tables, charts, or figures.
- The benchmark was verified by 194 human experts.
- Current systems struggle with long contexts and structured reasoning.
Paper Resources
Source Excerpt
Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages. Some questions also require comparing related reports. Reliable long-document understanding is therefore a prerequisite for using in compliance, clinical, financial, and engineering workflows, where decisions must be traceable to specific evidence pages and the cost of an unsupported answer is high -- yet
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.