
Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids
Quick Answer
Google Deepmind has launched Gemini Robotics 2, its most advanced vision-language-action model, capable of controlling various robots from tabletop arms to humanoids.
Quick Take
Additionally, the new Gemini Robotics ER 2 enhances embodied reasoning for higher-level control, replacing the previous ER 1.6 model.
Key Points
- Gemini Robotics 2 combines image recognition, language processing, and action control.
- The model can manage full-body movement and fine motor tasks for various robots.
- Developers can apply for early access to Gemini Robotics 2 via a waitlist.
- Gemini Robotics ER 2 serves as a higher-level control system for robots.
- ER 2 is now available in Google AI Studio, replacing the older ER 1.6 model.
📖 Reader Mode
~1 min readGoogle Deepmind has introduced Gemini Robotics 2, which it calls its most advanced vision-language-action (VLA) model yet. VLA models combine image recognition, language processing, and action control to help robots operate in physical environments. Deepmind says the model can control systems ranging from tabletop arms to full-body humanoid robots.
The company describes Gemini Robotics 2 as an "intelligence layer" for a new generation of adaptive robots. It can manage full-body movement, perform fine motor tasks, and coordinate multiple robots, according to Deepmind. Developers can apply for early access through the waitlist.
Google Deepmind also introduced Gemini Robotics ER 2, a model designed for "embodied reasoning." The term refers to understanding the physical world and deciding which actions to take based on that information. ER 2 acts as a higher-level control system for robots and replaces Gemini Robotics ER 1.6, released in April. The new model is available in Google AI Studio.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

