Post
Quick Answer
Thinking Machines Lab's Inkling-Small, a 276B Mixture-of-Experts model, surpasses its predecessor Inkling in multiple benchmarks, achieving SWE-bench 80.2% vs.
Quick Take
77.6%. With 42 decoder layers and a context window of 1M tokens, it operates efficiently on a single GPU, making it a significant advancement in model performance and resource optimization.
Key Points
- Inkling-Small achieved 80.2%, outperforming Inkling's 77.6%.
- Model utilizes 42 decoder layers with hybrid attention and 6 of 256 experts per token.
- Only ~4.3% of weights are active for any given token, enhancing efficiency.
- Fits on a single GPU with a BF16 checkpoint requiring 600 GB of VRAM.
- Training included a revised data mix and two weeks of agentic coding RL.
📖 Reader Mode
~1 min readThinking Machines Lab released Inkling-Small on July 30, 2026. It is a 276B total, 12B active Mixture-of-Experts model under Apache 2.0, roughly a quarter the size of Inkling. 1. The smaller model passed its teacher → SWE-bench Verified: 80.2% vs Inkling's 77.6% → Terminal-Bench 2.1: 64.7% vs 63.8% → HLE, text only: 31.6% vs 29.7% → ARC-AGI-2: 40.1% vs 36.5% Inkling-Small started training after Inkling. That let the team revise the pre-training data mix, distill from Inkling on-policy, then scale agentic coding RL for two more weeks. 2. Sparsity is doing the work → 42 decoder layers, hybrid local and global attention → 6 of 256 experts routed per token, plus 2 shared experts → ~4.3% of weights active on any given token → 1M token context window 3. It fits on one GPU → BF16 checkpoint: 600 GB aggregated VRAM (4x B300 or 8x H200) → NVFP4 checkpoint: 180 GB, W4A4 on 1x B300 (SM100+) or W4A16 on 2x H200 → Runtimes: SGLang, vLLM, TokenSpeed, Unsloth, Hugging Face Full analysis: marktechpost.com/2026/08/02/thi… Technical details: thinkingmachines.ai/news/inkling-s… Model weights: huggingface.co/thinkingmachin…
@thinkymachines— Originally published at x.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from WebSearch (Tavily)
See more →全球AI芯片峰会,9月上海见!
The 2026 Global AI Chip Summit will take place in Shanghai on September 22-23, focusing on the evolving AI chip landscape, including the shift from training to inference, the rise of diverse chip technologies, and the restructuring of industry competition. Notable speakers include experts from leading universities and companies, discussing advancements in AI chip architecture and applications.