PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs
Quick Answer
PatiGonit22K is a newly introduced Bengali dataset consisting of 22,441 mathematical word problems (MWPs), enhancing the original PatiGonit dataset.
Quick Take
This resource aims to improve natural language understanding and quantitative reasoning in Bengali, addressing the scarcity of large annotated datasets for low-resource languages.
Key Points
- Dataset includes 22,441 problems, enhancing mathematical reasoning resources for Bengali.
- Problems range from simple to multi-operation equations, providing varied difficulty levels.
- Each problem is translated, annotated, and culturally adapted for consistency and correctness.
- PatiGonit22K addresses the lack of large-scale annotated datasets in low-resource languages.
- Aims to support future research in educational NLP applications for Bengali.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning. Despite recent progress in high resource languages, Bengali remains underexplored due to the limited availability of large scale annotated datasets. In this work, we introduce PatiGonit22K, an expanded Bengali MWP dataset containing 22,441 problems, developed by extending the original PatiGonit dataset with a substantially larger collection of complex mathematical problems. The dataset includes both simple and multi operation equations, providing a balanced benchmark for evaluating mathematical reasoning across different difficulty levels. Each problem is carefully translated, annotated, culturally adapted, and verified to ensure linguistic consistency and mathematical correctness. By increasing both the scale and complexity of Bengali MWPs, PatiGonit22K provides a more comprehensive resource for future research on mathematical reasoning and educational NLP applications in low resource languages.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.22859 [cs.CL] |
| (or arXiv:2607.22859v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22859 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Swastika Kundu [view email]
[v1]
Fri, 24 Jul 2026 19:01:17 UTC (313 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.