Towards Structuring an Arabic-English Machine-Readable Dictionary Using Parsing Expression Grammars
Quick Answer
This paper proposes a method to convert the Arabic-English Al-Mawrid dictionary into a machine-readable format using parsing expression grammars.
Quick Take
The approach structures dictionary entries into hierarchical formats, enhancing their usability for natural language processing applications despite the lack of standardization in Arabic dictionaries.
Key Points
- The method structures dictionary entries into hierarchical formats for better machine processing.
- Parsing expression grammars were utilized to implement the parser for the dictionary.
- Each dictionary entry includes subentries with defining phrases and translation equivalences.
- The study shows potential for automatic or semi-automatic structuring of Arabic dictionaries.
- Lack of microstructure standardization in Arabic dictionaries is addressed through this method.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Dictionaries are rich sources of lexical information about words that is required for many applications of natural language processing and human language technology. However, publishers prepare printed dictionaries for human usage not for machine processing. This paper presented a method to structure partly a machine-readable version of the Arabic-English Al-Mawrid dictionary. The method converted the entries of Al-Mawrid from a stream of words and punctuation marks into hierarchical structures. The hierarchical structure expresses the components of each dictionary entry in explicit format. A dictionary entry is composed of subentries and each subentry consists of defining phrases, domain labels, cross-references, and translation equivalences. We designed the proposed method as cascaded steps where parsing is the main step. We implemented the parser using the parsing expression grammars formalism. In conclusion, although Arabic dictionaries do not have microstructure standardization, this study demonstrated that it is possible to structure them automatically or semi-automatically with plausible accuracy after inducing their microstructure.
| Comments: | 14 pages, 6 figures, 7 tables. The final publication is available at this https URL. Published in International Journal of Computational Linguistics Research (IJCLR), DLINE, March 2014, Vol 5, Issue 1, pp 1-13 |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2606.25231 [cs.CL] |
| (or arXiv:2606.25231v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.25231 arXiv-issued DOI via DataCite (pending registration) |
|
| Journal reference: | International Journal of Computational Linguistics Research (IJCLR), 5(1), pp 1-13, March 2014, DLINE Publisher |
Submission history
From: Diaa Fayed [view email]
[v1]
Tue, 23 Jun 2026 23:17:51 UTC (1,029 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.