MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities
Quick Answer
MioFFAn is an open-source annotation framework designed for Formula Formalization, enhancing the MioGatto architecture with customizable features for scientific equations.
Quick Take
It integrates for partial automation, allowing researchers to refine automation strategies through a modular approach and evaluate them with standard NLP metrics.
Key Points
- MioFFAn facilitates rapid annotation for translating mathematical expressions into executable code.
- The framework allows customization of taxonomies and properties for scientific symbols.
- It supports partial automation using Large Language Models for iterative refinement.
- Researchers can evaluate competing strategies with standard NLP metrics.
- The framework addresses the scarcity of high-quality datasets in technical scientific domains.
DeepSignal Analysis
What happened
MioFFAn is an open-source framework aimed at enhancing Formula Formalization by integrating Large Language Models for partial automation. It builds on the MioGatto architecture, allowing users to customize features for scientific equations. The framework supports the iterative refinement of automation strategies through a modular approach and evaluation using standard NLP metrics.
Key evidence
- MioFFAn is designed to facilitate rapid annotation for Formula Formalization, addressing the lack of high-quality datasets in technical scientific domains.
- The framework allows users to configure custom taxonomies and properties for identified symbols, making it adaptable to various specialized scientific fields.
- A preliminary evaluation of MioFFAn demonstrates the efficacy of its human-in-the-loop approach for refining automation capabilities.
Why it matters
The development of MioFFAn addresses a significant gap in the automation of translating mathematical expressions into executable code, which is crucial for advancing computational research. By enabling customizable annotation and integrating LLMs, it provides a flexible tool for researchers in diverse scientific fields. This could lead to more efficient workflows and improved accuracy in scientific computations.
Paper Resources
📖 Reader Mode
~2 min readAbstract:The automatic translation of mathematical expressions in scientific literature into executable symbolic code (a process we refer to as Formula Formalization) is hindered by a severe scarcity of high-quality, ground-truth datasets specialized for technical scientific domains. In this paper, we present MioFFAn, an open-source, document-centric, and customizable framework designed to facilitate rapid annotation for this task. Building upon the MioGatto architecture, we extend existing features to overcome structural limitations and pivot its scope by introducing specific functionalities for Formula Formalization, such as selection of equations of interest and aided symbolic code specification. By allowing users to configure custom taxonomies and properties for identified symbols, and compatible symbolic operators, we ensure the framework is adaptable to diverse specialized scientific fields. Furthermore, MioFFAn is designed to incorporate partial automation via Large Language Models. By defining a modular set of automated sub-tasks with strict output formats, we enable researchers to iteratively refine automation capabilities and evaluate competing strategies using standard NLP metrics. We specify the current automation methodology and perform a preliminary evaluation that demonstrates to efficacy of this human-in-the-loop approach.
| Comments: | Presented in the 3rd International Workshop on Natural Scientific Language Processing (NSLP 2026), co-located at LREC2026 |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG); Software Engineering (cs.SE) |
| Cite as: | arXiv:2607.22552 [cs.CL] |
| (or arXiv:2607.22552v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22552 arXiv-issued DOI via DataCite |
|
| Journal reference: | Proceedings of the 3rd Int. Workshop on Natural Scientific Language Processing (NSLP 2026) at LREC 2026, pages 206-217 |
Submission history
From: Nicolas Sibuet [view email]
[v1]
Fri, 15 May 2026 13:01:27 UTC (990 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.