
Poolside's Laguna S 2.1 is a small open-weight coding model that punches well above its size
Quick Answer
Poolside's Laguna S 2.1, a compact coding model, achieves 70.2% on Terminal-Bench 2.1, outperforming larger models like DeepSeek-V4-Pro-Max.
Quick Take
Its thinking mode significantly enhances performance, with a notable cost of just $0.088 for solving Erdős Problem 397, a math challenge unsolved since 1975.
Key Points
- Laguna S 2.1 scores 40.4% on Datacurve's DeepSWE benchmark.
- The model's thinking mode boosts performance, with a score drop to 60.4% without it.
- Poolside emphasizes persistence and verification over merely scaling model size.
- Laguna S 2.1 can run locally or via hosted services like Hugging Face.
- The model was trained using 4,096 Nvidia H200 GPUs over nine weeks.
DeepSignal Analysis
What happened
Poolside's Laguna S 2.1, released in July 2026, demonstrates strong performance on various benchmarks, achieving 70.2% on Terminal-Bench 2.1. The model's thinking mode significantly enhances its capabilities, allowing it to solve complex problems like Erdős Problem 397 at a low cost. The model is available under the OpenMDW 1.1 license, enabling broad usage and modification.
Key evidence
- Laguna S 2.1 scored 70.2% on Terminal-Bench 2.1, outperforming larger models such as DeepSeek-V4-Pro-Max and Nemotron 3 Ultra.
- The model's thinking mode improved its Terminal-Bench score from 60.4% to 70.2%, indicating a significant performance gap between its two operational modes.
- Poolside's training for Laguna S 2.1 involved 409,000 environments, including 83,000 for terminal tasks, contributing to its enhanced performance.
Why it matters
The performance of Laguna S 2.1 challenges the notion that larger models are inherently better, suggesting that improvements in model behavior can lead to significant gains. Its ability to solve long-standing mathematical problems indicates potential for practical applications in various fields. The model's open licensing may encourage wider adoption and innovation in AI development.
Source Excerpt
Poolside has released Laguna S 2. 1, its third coding model in three months. Rather than rely on raw scale, the company trained it to keep checking its work, revise failed approaches, and avoid giving up too soon during long agentic sessions. The compact model beats several much larger rivals in benchmarks. Poolside says it also solved a math problem that had been open since 1975 for under 10 cents.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

