JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents
Quick Answer
JetBrains has launched Mellum2.1, a 12B mixture-of-experts model featuring 2.5B active parameters.
Quick Take
This model achieved a significant Verified score increase from 2.0 to 47.0 through reinforcement learning applied in real repositories, enhancing coding agent capabilities.
Key Points
- Mellum2.1 is an Apache 2.0 licensed model with 12 billion parameters.
- The model utilizes a mixture-of-experts architecture with 2.5 billion active parameters.
- Reinforcement learning in real repositories improved its SWE-bench score significantly.
- The SWE-bench Verified score rose from 2.0 to 47.0.
- Mellum2.1 aims to enhance the performance of coding agents.
Article Excerpt
From source RSS / original summaryJetBrains released Mellum2. 1, an Apache 2. 0, 12B mixture-of-experts thinking model with 2. 5B active parameters. RL in real repositories lifted its Verified score from 2. 0 to 47. 0. The post JetBrains Releases Mellum2. 1: A 12B MoE Open Model for Coding Agents appeared first on MarkTechPost.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from MarkTechPost
See more →Meet Flash-KMeans: An IO-Aware, Exact K-Means That Runs Over 200× Faster Than FAISS on GPUs
Flash-KMeans is an open-source, IO-aware k-means implementation that operates over 200× faster than FAISS on NVIDIA H200 GPUs. It achieves 17.9× end-to-end and 33× speedup over cuML by optimizing distance calculations and updating mechanisms without approximating results. This advancement significantly enhances performance for data scientists and machine learning practitioners.