RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluation
Quick Answer
RusFinChain is the first Russian-language benchmark for verifiable Chain-of-Thought reasoning in finance, featuring 5,280 examples across 17 domains.
Quick Take
Evaluation of 8 open-weight shows a Hard F1 score of ~0.65 for step alignment, but only ~29% of final answers are correct, highlighting a significant reasoning gap.
Key Points
- RusFinChain includes 5,280 parameterized examples from executable Python templates.
- Enhanced metrics like Fuzzy Numeric Alignment show better correlation with answer correctness.
- Models achieved ~0.65 Hard F1 for step alignment but only ~29% final answer accuracy.
- Dataset and evaluation framework released to support Russian-speaking financial AI development.
- Evaluation involved 8 open-weight LLMs generating 8,100 responses.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Multi-step symbolic reasoning is essential for robust financial analysis, yet most benchmarks neglect intermediate reasoning steps. FINCHAIN introduced verifiable Chain-of-Thought (CoT) evaluation but is limited to English. FINESSE-Bench includes a Russian block but relies on multiple-choice questions without step-level supervision. We present RusFinChain, the first Russian-language symbolic benchmark for verifiable CoT reasoning in finance. It spans 17 domains, 172 topics, and comprises 5,280 parameterized examples from executable Python templates, ensuring contamination-free evaluation. Each example includes a gold-standard reasoning chain with intermediate numeric values for automatic verification. We also introduce enhanced metrics: Fuzzy Numeric Alignment and Soft-Attention Alignment. We evaluate 8 open-weight LLMs on a stratified sample, generating 8,100 responses. Results reveal a substantial reasoning gap: models achieve Hard F1 of ~0.65 for step alignment, but only ~29% of final answers are correct. Our fuzzy and soft metrics show stronger correlation with final-answer correctness (Spearman rho approx 0.48) than the original ChainEval (rho approx 0.38-0.46), demonstrating superior diagnostic power. We release dataset, code, and evaluation framework to foster verifiable financial AI for the Russian-speaking community.
| Comments: | Preprint |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.01388 [cs.CL] |
| (or arXiv:2607.01388v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01388 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mullosharaf Arabov Am [view email]
[v1]
Wed, 1 Jul 2026 18:48:05 UTC (38 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.