
Thinking Machines bets on efficiency over size with its second model, Inkling Small
Quick Answer
Thinking Machines has launched Inkling Small, an efficient reasoning model scoring 40 on the Intelligence Index with 276 billion parameters.
Quick Take
It outperforms its predecessor in coding tests while being more token-efficient, averaging 24K output tokens per task, and offers a 256K-token context window for diverse inputs.
Key Points
- Inkling Small has 276 billion parameters and scores 40 on the Intelligence Index.
- It outperforms Inkling in coding tests like (32% vs. 30%).
- The model is highly token-efficient, averaging 24K output tokens per task.
- It supports text, image, and speech inputs with a 256K-token context window.
- Weights are available on Hugging Face for user fine-tuning.
📖 Reader Mode
~1 min readThinking Machines, the AI lab from former OpenAI CTO Mira Murati, has released Inkling Small. According to Artificial Analysis, the open-weights reasoning model scores 40 on the Intelligence Index, one point below Inkling (41), with less than a third of the parameters (276 billion total, 12 billion active). AA says no open model of equal or smaller size scores higher.
Inkling Small beats its bigger sibling on several coding and reasoning tests, including Humanity's Last Exam (32% vs. 30%) and GPQA Diamond (89% vs. 87%). It falls behind on agent-based tasks and factual knowledge but is far more token-efficient, averaging 24K output tokens per task compared to 45K for Deepseek V4 Flash and 78K for GPT-5.4 mini.

The model handles text, image, and speech inputs, has a 256K-token context window, and ships under Apache 2.0. Weights are on Hugging Face, and users can fine-tune it in the browser via Tinker Playground. Thinking Machines positions its models as a foundation for fine-tuning with users' own data. Some see this as the next frontier in AI.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

