Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
Quick Answer
This paper shows that FinMMEval 2026 Task 2 assesses multilingual financial short-answer question answering, featuring 256 test items across five languages.
Quick Take
Systems submitted concise answers in JSONL format, with top performers closely ranked by ROUGE-1 F1 scores, highlighting advancements in and cross-lingual evidence handling.
Key Points
- Task features 256 items, split between easy and expert tiers.
- Final leaderboard includes 12 ranked submissions with close ROUGE-1 F1 scores.
- Participating languages include English, Chinese, Japanese, Spanish, and Greek.
- Gold answers were withheld to ensure unbiased evaluation.
- Systems utilized strategies like structured prompting and answer compression.
Paper Resources
Source Excerpt
FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and news in English, Chinese, Japanese, Spanish, and Greek. Participating systems submit one concise answer per item in JSONL format. The final-test set contains 256 items, split evenly between easy and expert tiers; each tier contains four question templates instantiated over 32 company-report groups. Gold answers were wit
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.