Marktechpost AI on X: "Thinking Machines Lab released Inkling-Small on July 30, 2026. It is a 276B total, 12B active Mixture-of-Experts model under Apache 2.0, roughly a quarter the size of Inkling. 1. The smaller model passed its teacher → SWE-bench Verified: 80.2% vs Inkling's 77.6% → Terminal-Ben
Quick Answer
Thinking Machines Lab launched Inkling-Small, a 276B total, 12B active Mixture-of-Experts model, achieving benchmark scores of 80.2% on SWE-bench, outperforming its predecessor Inkling.
Quick Take
The model's sparsity allows it to fit on a single GPU, making it efficient for deployment.
Key Points
- Inkling-Small is a 276B total, 12B active model under Apache 2.0.
- Achieved 80.2% on , surpassing Inkling's 77.6%.
- Utilizes 42 decoder layers with 4.3% weights active per token.
- Fits on a single GPU with 600 GB aggregated VRAM for BF16 checkpoint.
- Started training after Inkling, allowing for improved pre-training data mix.
📖 Reader Mode
~1 min readThinking Machines Lab released Inkling-Small on July 30, 2026. It is a 276B total, 12B active Mixture-of-Experts model under Apache 2.0, roughly a quarter the size of Inkling. 1. The smaller model passed its teacher → SWE-bench Verified: 80.2% vs Inkling's 77.6% → Terminal-Bench 2.1: 64.7% vs 63.8% → HLE, text only: 31.6% vs 29.7% → ARC-AGI-2: 40.1% vs 36.5% Inkling-Small started training after Inkling. That let the team revise the pre-training data mix, distill from Inkling on-policy, then scale agentic coding RL for two more weeks. 2. Sparsity is doing the work → 42 decoder layers, hybrid local and global attention → 6 of 256 experts routed per token, plus 2 shared experts → ~4.3% of weights active on any given token → 1M token context window 3. It fits on one GPU → BF16 checkpoint: 600 GB aggregated VRAM (4x B300 or 8x H200) → NVFP4 checkpoint: 180 GB, W4A4 on 1x B300 (SM100+) or W4A16 on 2x H200 → Runtimes: SGLang, vLLM, TokenSpeed, Unsloth, Hugging Face Full analysis: marktechpost.com/2026/08/02/thi… Technical details: thinkingmachines.ai/news/inkling-s… Model weights: huggingface.co/thinkingmachin…
@thinkymachines— Originally published at x.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from WebSearch (Tavily)
See more →全球AI芯片峰会,9月上海见!
The 2026 Global AI Chip Summit will take place in Shanghai on September 22-23, focusing on the evolving AI chip landscape, including the shift from training to inference, the rise of diverse chip technologies, and the restructuring of industry competition. Notable speakers include experts from leading universities and companies, discussing advancements in AI chip architecture and applications.


