Google DeepMind Ships Three Physical AI Models For Whole Body Control
A major leap in robotic intelligence combining whole-body control, dexterous manipulation, and multi-robot collaboration via Gemini Robotics 2.

Stock photo for illustration only, not from the actual event
- Gemini Robotics 2 controls humanoids from feet to fingertips and dual arms
- ER 2 acts as a high-level brain built upon Gemini 3.5 Flash
- On-Device 2 is optimized to run locally on robotic hardware
- Enables seamless collaboration between different robot embodiments
Most robots today remain constrained by pre-programmed routines or tele-operation designed for narrow, repetitive tasks, struggling to adapt in unpredictable environments or transfer skills across robot bodies. Gemini Robotics 2 directly targets all three limitations at once by converting vision and language inputs into precise motor control. It drives full humanoids from feet to fingertips, operates bi-arm robots, and handles dexterous manipulation using multi-finger hands as well as parallel grippers.
Acting as the high-level brain is Gemini Robotics ER 2, an embodied reasoning model built on vision-language technology. According to its model card, ER 2 is based on Gemini 3.5 Flash, accepting interleaved text, image, video, and audio with a context window of up to 128k while outputting text up to 64K tokens to communicate with humans and plan multi-minute tasks.

Stock photo for illustration only, not from the actual event
System design relies on a clear division of labor where ER 2 plans and tracks tasks, subsequently delegating motor execution to a VLA declared as a tool. Developers can register low-level control interfaces—such as VLA models or navigation APIs—as callable tools, streaming multimodal video, audio, or text directly into the framework.
Regarding task progress tracking, ER 2 achieves 57.4% accuracy when classifying video frames into five progress levels. It reaches 91.3% accuracy with a 0.96-second mean absolute distance when determining the exact frame for critical events, such as knowing precisely when to stop pouring coffee into a cup—competing closely with much larger model categories at four times the execution speed, as reported by Google DeepMind.
"Google DeepMind is direct about the remaining gap. It states that its robots have more to advance in movement speed."
The introduction of the Gemini Robotics 2 suite illustrates a pivotal shift in physical AI toward closely coupling cognitive reasoning with physical actuation. By decoupling high-level task planning from low-level motor execution, these models grant robots unprecedented adaptability, mirroring how humans conceptualize goals before engaging muscle movements to overcome traditional rigid programming boundaries.

Stock photo for illustration only, not from the actual event
Furthermore, Gemini Robotics On-Device 2 is an efficient VLA optimized to execute locally on robots, bypassing network latency and connectivity constraints. Built on Gemini Robotics 1.5 and Google's on-device Gemma models, it processes text, images, and robot proprioception as numerical values to output direct mechanical actions.
The framework also pioneers multi-robot collaboration by allowing distinct robot embodiments—such as Apptronik's Apollo 2 paired with a Franka F3 Duo—to communicate through a shared semantic understanding and seamlessly hand off subtasks depending on terrain and operational strengths.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment