
Inkling Small from Thinking Machines is now available on AI Gateway
Quick Answer
Inkling Small from Thinking Machines is now available on AI Gateway, offering performance similar to larger models at a quarter of the size.
Quick Take
It excels in reasoning, coding, and image processing while allowing users to control the trade-off between quality, cost, and latency. This model is also compatible with Zero Data Retention for enhanced privacy.
Key Points
- Inkling Small achieves comparable performance to larger models with reduced compute requirements.
- Supports advanced reasoning, coding, and image processing tasks effectively.
- Users can adjust reasoning effort to balance quality, cost, and latency.
- Compatible with Zero Data Retention for enhanced privacy settings.
- AI Gateway offers no markup on provider pricing and no platform fees.
DeepSignal Analysis
What happened
Thinking Machines has released Inkling Small on AI Gateway, a model that offers performance similar to its larger counterpart while being significantly smaller in size. This model is designed for various tasks, including reasoning, coding, and image processing, and allows users to adjust the balance between quality, cost, and latency. Additionally, it supports Zero Data Retention for enhanced privacy.
Key evidence
- Inkling Small achieves performance comparable to the larger Inkling model while being about a quarter of its size, using less compute per task.
- The model is capable of performing reasoning over audio and images, and it can execute tasks such as cropping and inspecting images programmatically.
- Inkling Small is compatible with Zero Data Retention, allowing users to ensure that prompts and responses are deleted after each request.
Why it matters
The introduction of Inkling Small could provide a more accessible option for users needing AI capabilities without the resource demands of larger models. Its ability to balance quality, cost, and latency makes it suitable for a range of applications, particularly in coding and image processing. Moreover, the Zero Data Retention feature addresses growing concerns about data privacy, making it a more appealing choice for organizations focused on compliance and security.
📖 Reader Mode
~2 min readInkling Small from Thinking Machines is now available on AI Gateway.
Inkling Small reaches performance comparable to the larger Inkling model at about a quarter of the size, using much less compute per task. It is a broad generalist with native reasoning over audio and images, and it holds up well on reasoning, agentic coding, and tool use. Controllable thinking effort lets you trade quality against cost and latency, from minimal to maximum reasoning.
For visual tasks, it can crop, zoom, and inspect images programmatically, which helps on documents and charts where the relevant detail is small.
To use Inkling, set model to thinkingmachines/inkling-small in the AI SDK:
import { streamText } from 'ai';
const result = streamText({
model: 'thinkingmachines/inkling-small',
prompt: 'Summarize this report and list the key risks.',
});
Inkling-Small is compatible with Zero Data Retention. Turn it on team-wide from the dashboard, or per request with zeroDataRetention: true, and AI Gateway routes only to providers that delete prompts and responses after each request.
Inkling-Small is also a cost-efficient choice for coding and tool-use workflows. Run vercel ai-gateway coding-agents setup to connect your coding agents to AI Gateway, then select thinkingmachines/inkling-small in the agent's model configuration. See the coding agents guide.
AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Try Inkling Small in the model playground.
AI Gateway: Track top AI models by usage
The AI Gateway model leaderboard tracks the most popular models over time, ranking them by the total volume of tokens processed across all Gateway traffic.
View the leaderboard
— Originally published at vercel.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Vercel AI
See more →
The Agent Stack
The Agent Stack by Vercel AI provides essential building blocks for creating production-grade agents, enabling seamless integration across multiple AI models and secure operations. It features components like AI Gateway for model routing, Workflow SDK for durable execution, and Vercel Connect for scoped access, streamlining agent development and deployment across various platforms.

