
Mistral's open-source Leanstral 1.5 aces formal math benchmarks and catches real bugs in code
Quick Answer
Mistral AI has launched Leanstral 1.5, an open-source model for formal verification in Lean 4, which excelled in formal math benchmarks and identified five previously unknown bugs across 57 open-source repositories.
Key Points
- Leanstral 1.5 is designed for formal verification using Lean 4.
- The model excelled in formal math benchmarks.
- It discovered five previously unknown bugs in open-source code.
- The bugs were found while scanning 57 repositories.
- This release enhances reliability in software development.
📖 Reader Mode
~1 min readMistral AI released Leanstral 1.5, a free open-source model (Apache 2.0 license) built for formal verification in the Lean 4 programming language. Lean 4 is designed to formally verify mathematical proofs and software correctness.
Mistral says the model hits 100 percent on miniF2F, a formal math benchmark covering problems from high school level up to math olympiad difficulty. On PutnamBench, which includes 672 problems from the Putnam math competition, it solves 587. On the algebra benchmarks FATE-H and FATE-X, which test master's and doctoral-level tasks in areas like group theory and ring theory, it scores top results of 87 and 34 percent.

The model was trained mainly for math, but Mistral says it also performs well at code verification. In a hands-on test, it scanned 57 open-source repositories and caught five previously unknown bugs, including an overflow bug in the Rust library varinteger. The model is available through Hugging Face and a free API. Training involved mid-training, supervised fine-tuning, and reinforcement learning.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

