
Google's Gemini 4 "Carbon" model is reportedly matching Anthropic's Opus 5.5 coding performance
Quick Answer
Google's Gemini 4 'Carbon' model is reportedly outperforming Argon in coding tasks and is comparable to Anthropic's Opus 5.5.
Quick Take
While still undergoing testing, Carbon is part of a rapid iteration process that showcases Google's advancements in AI model development.
Key Points
- Gemini 4 'Carbon' is outperforming Argon in coding tasks.
- Carbon is compared to Anthropic's Opus 5.5, pending further testing.
- Google's models are mapped to specific roles for optimized performance.
- Rapid updates indicate improved AI-driven model development at Google.
- No official launch date for Gemini 4 has been announced yet.
DeepSignal Analysis
What happened
Google is testing multiple variants of its Gemini 4 model, including Carbon, which reportedly outperforms Argon in coding tasks and is comparable to Anthropic's Opus 5.5. The company is iterating rapidly on these models, with Carbon being deployed on its internal coding platform, Jetski. However, further testing is needed to confirm its performance.
Key evidence
- Carbon has been deployed on Google's internal coding platform Jetski and is said to outperform Argon in programming tasks.
- An employee compared Carbon to Anthropic's Opus 5.5, indicating it may match or exceed that model's coding performance.
- Google is preparing for the Gemini 4 launch, with updates to its apps and features like a new 'Automatic' mode and adjustable reasoning intensity.
Why it matters
The development of the Gemini 4 models, particularly Carbon, highlights Google's ongoing efforts to enhance its AI capabilities. By comparing Carbon to Anthropic's Opus 5.5, Google positions itself competitively in the AI coding space. The rapid iteration process suggests advancements in AI model development, which could lead to improved performance and functionality in various applications.
📖 Reader Mode
~2 min readGemini 4 "Argon" hasn't even shipped yet, and there are already rumors about the next performance update.
According to documents, screenshots, and internal chats seen by Business Insider, Google is testing several Gemini 4 variants called Argon, Barium, and Carbon. Carbon was deployed on Google's internal coding platform Jetski over the past few days and is said to beat Argon mainly on programming tasks.
One employee compared Carbon to Anthropic's strongest coding model, Opus 5.5, though it still needs more testing. Early Argon versions, by contrast, reminded another employee of the older Opus 5 on some coding tasks, even though internal feedback was positive overall.
Argon is Google's frontier reasoning model
Early rumors pegged Argon as a Flash model optimized for speed over peak performance. The announcement of Google's Gemini agent for Google Workspace puts that to rest. Google maps its models to specific roles: "Argon for frontier reasoning, Flash for speed and volume, Omni for generative media, and Gemma for lightweight, open-weights edge workloads," the company writes.
So Argon is definitely Google's most capable reasoning model, comparable to Anthropic's Opus or OpenAI's Astra. Carbon and Barium are likely checkpoints or updates within the Argon family, not separate tiers. One employee internally called Carbon "Gemini pro next model," suggesting it would ship under the Argon name. That fits with internal documents showing the now-unveiled Argon was previously called "Barium-B." Whether Carbon ships as an Argon update or a standalone model remains unclear.
The quick update pace, similar to what Google has done with the Flash models, suggests Google has gotten better at using AI to build new models. Google Deepmind employee Vedant Misra responded to the Business Insider report on X by writing, "Have you heard of recursive self improvement," referring to using AI to build better AI. OpenAI and Anthropic have reported similar progress.

Google is already prepping for the Gemini 4 launch
There's still no official Gemini 4 launch date yet, but Google is already retooling its apps. The Gemini app now shows a new "Automatic" mode for the 3-series models and adjustable reasoning intensity from low to high. Google AI Studio has surfaced a new "Ultra" mode promising "advanced skills and tools." Google's Logan Kilpatrick confirmed the team is working on getting the most out of Argon.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

