Whole-body robot control
Controls full humanoid robots from feet to fingertips, translating vision and language into motor actions for walking, crouching, reaching, and manipulation.
Gemini Robotics 2 by Google DeepMind is a physical AI model family for robots, with whole-body control, dexterous manipulation, embodied reasoning and on-device operation.
Gemini Robotics 2 is Google DeepMind’s physical AI model family for robots. The site positions it as an intelligence layer for adaptable robots, combining whole-body control, dexterous manipulation, embodied reasoning, and on-device operation.
The product page presents three related models: Gemini Robotics 2, a vision-language-action model for motor control; Gemini Robotics ER 2, a vision-language model for high-level planning and human communication; and Gemini Robotics On-Device 2, an efficient VLA model that runs locally on robot hardware. Together they are aimed at robots that need to perceive, reason, act, and coordinate across a range of embodiments.
Controls full humanoid robots from feet to fingertips, translating vision and language into motor actions for walking, crouching, reaching, and manipulation.
Adds a higher level of physical dexterity for both five-finger hands and standard grippers, including delicate and packing-style tasks.
Lets robots reason about tasks that take minutes and involve many steps, then coordinate execution with the VLA model while tracking progress.
Enables multiple robots to communicate and delegate work to complete a shared mission when one robot is not enough.
Runs locally on robotic devices and is designed for fast adaptation to new robot embodiments with only a few hours of data.
A humanoid robot receives a natural-language instruction, walks to a location, bends or crouches as needed, and places an object in a target bin or shelf.
A robot uses a five-finger hand or a parallel gripper to perform detailed manipulation such as knot-tying, ziplock sealing, or tight packing.
A robot plans a longer workflow, tracks whether each step succeeded, retries when needed, and keeps going until the task is complete.
Multiple robots coordinate in the same environment, communicate their strengths, and divide work to finish a shared mission.
A local robot deployment adapts a model to a new embodiment without relying on network connectivity or cloud latency.
Gemini Robotics ER 2 is the embodied reasoning model. It understands the physical world, plans multi-step tasks, communicates with humans, and hands off motor execution to a lower-level VLA model.
The source says Gemini Robotics 2 is available as a vision-language-action model for controlling robots, Gemini Robotics ER 2 is available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform, and the VLA and On-Device models are available to early-access partners.
The pages describe a workflow where ER 2 handles reasoning, planning, progress tracking, and coordination, while the VLA model performs motor control on the robot.
The on-device model is optimized to run locally on robotic devices and is designed for fast adaptation to new robot embodiments with a few hours of data.
The source notes that whole-body control and medium-to-high success rates are demonstrated for whole-body and gripper-based dexterous tasks, while multi-finger dexterous manipulation remains challenging.