
Coinbase joins the rush to Chinese AI models as Western labs face a pricing stress test
Quick Answer
Coinbase is adopting Chinese AI models like GLM 5.2 and Kimi 2.7, utilizing an automated routing system that optimizes model selection based on task and cost.
Quick Take
This shift has halved their AI spending while increasing token usage, with caching improvements boosting hit rates from 5% to 60%.
Key Points
- Coinbase's AI spending has been reduced by 50%.
- Token usage at Coinbase continues to rise despite spending cuts.
- Automated routing selects the best AI model based on task and price.
- Hit rate for model requests improved from 5% to 60% due to better caching.
- Chinese models GLM 5.2 and Kimi 2.7 are now preferred by Coinbase.
📖 Reader Mode
~2 min readCoinbase CEO Brian Armstrong has moved his company to cheap Chinese AI models. The company is using more tokens than ever but paying half what it used to.
Coinbase now runs on models like GLM 5.2 and Kimi 2.7, according to Armstrong. Developers can still pick whatever model they want, but 91 percent never hit their old usage limits anyway.
The CEO of startup Lindy made the same move to Deepseek v4 recently. Snowflake is testing Chinese models too as cheaper alternatives to OpenAI and Anthropic. That puts real pricing pressure on Western AI labs and adds risk right as some are eyeing IPOs. It's a stress test for the growth numbers they need to hit to justify the money they've raised.
Coinbase also runs an automatic routing system that picks the best model for each request based on task, price, and caching potential. Better caching alone pushed the hit rate from 5 to 60 percent. Developers are told to keep context lean and start fresh sessions for new tasks, a strategy that falls under the broader umbrella of context engineering.

Tokenmaxxing meets accountability
Coinbase also makes each developer's usage visible without capping it. That echoes the tokenmaxxing trend where employees at Amazon and Meta got kudos for burning through tokens with no need to justify the results.
But Coinbase adds one rule that breaks that cycle. "The more you spend on AI, the more impact we expect," Armstrong says. These moves cut Coinbase's AI spending in half even as token usage keeps climbing, according to Armstrong.
More companies are leaning into this kind of optimization. We cover the rise of the token economy in our Frontier Radar #3.
Reportedly, a price war between OpenAI and Anthropic is brewing, too. OpenAI's GPT-5.6-Sol costs the same as GPT-5.5 but is supposed to be more token-efficient than Claude Fable and Mythos. OpenAI is also offering two weaker 5.6 variants at much lower prices.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

