
Ling 3.0 Flash is now available on AI Gateway
Quick Answer
Ling 3.0 Flash, a 124B parameter Mixture-of-Experts model from Ant Group, is now available for free on AI Gateway for three weeks.
Quick Take
It features a 256K token context window and is optimized for high-frequency workflows and multi-step agent interactions.
Key Points
- Ling 3.0 Flash operates with 5.1B active parameters per token.
- The model supports both thinking and non-thinking modes for flexible usage.
- AI Gateway offers a unified API for model calls and usage tracking.
- No platform fees are charged on inference, including BYOK requests.
- Try Ling 3.0 Flash in the model playground for hands-on experience.
📖 Reader Mode
~1 min readLing 3.0 Flash from Ant Group is now available on AI Gateway.
The model is free to use for the next three weeks, through August 3rd.
Ling 3.0 Flash is a Mixture-of-Experts model with 124B total parameters and about 5.1B active per token. It has a 256K token context window and runs in thinking and non-thinking modes.
Ling 3.0 Flash is built for token-efficient agentic inference at production scale, doing more work within tighter token, latency, and cost budgets across multi-step agent runs. The model targets high-frequency agentic workflows, coding agents, document work, and long-context multi-turn interactions.
To use Ling 3.0 Flash, set model to inclusionai/ling-3.0-flash-free in the AI SDK:
import { streamText } from 'ai';
const result = streamText({
model: 'inclusionai/ling-3.0-flash-free',
prompt: 'Triage the open issues in this repo and group them by theme.',
});
AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, routing rules, and more.
AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests.
Try Ling 3.0 Flash in the model playground.
— Originally published at vercel.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Vercel AI
See more →
The Agent Stack
The Agent Stack by Vercel AI provides essential building blocks for creating production-grade agents, enabling seamless integration across multiple AI models and secure operations. It features components like AI Gateway for model routing, Workflow SDK for durable execution, and Vercel Connect for scoped access, streamlining agent development and deployment across various platforms.

