Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
Quick Answer
This paper shows that FinMMEval 2026 Task 2 assesses multilingual financial short-answer question answering, featuring 256 test items across five languages.
Quick Take
Systems submitted concise answers in JSONL format, with top performers closely ranked by ROUGE-1 F1 scores, highlighting advancements in and cross-lingual evidence handling.
Key Points
- Task features 256 items, split between easy and expert tiers.
- Final leaderboard includes 12 ranked submissions with close ROUGE-1 F1 scores.
- Participating languages include English, Chinese, Japanese, Spanish, and Greek.
- Gold answers were withheld to ensure unbiased evaluation.
- Systems utilized strategies like structured prompting and answer compression.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Zhuohan Xie, Xueqing Peng, Georgi Georgiev, Dimitar Dimitrov, Yuyang Dai, Rania Elbadry, Vanshikaa Jani, Lingfei Qian, Fan Zhang, Jimin Huang, Jiahui Geng, Yankai Chen, Ye Yuan, Haolun Wu, Yuxia Wang, Ivan Koychev, Veselin Stoyanov, Mingzi Song, Yu Chen, Xue Liu, Preslav Nakov
Abstract:FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tier contains four question templates instantiated over 32 company-report groups. Gold answers were withheld during submission, and systems were ranked by macro-averaged item-level ROUGE-1 F1 against organizer-held reference answers. The final leaderboard includes 12 ranked submissions. The strongest systems are closely clustered, with the top four separated by less than one percentage point in ROUGE-1 F1. The submitted system papers document retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.
| Comments: | 9 pages. Task overview paper for CLEF 2026 Working Notes (CEUR Workshop Proceedings) |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE) |
| Cite as: | arXiv:2607.19867 [cs.CL] |
| (or arXiv:2607.19867v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.19867 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhuohan Xie [view email]
[v1]
Wed, 22 Jul 2026 07:54:17 UTC (105 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.