Gemini Robotics 2

Revendiquer

Gemini Robotics 2 by Google DeepMind is a physical AI model family for robots, with whole-body control, dexterous manipulation, embodied reasoning and on-device operation.

Gemini Robotics 2 preview

Overview

Gemini Robotics 2 is Google DeepMind’s physical AI model family for robots. The site positions it as an intelligence layer for adaptable robots, combining whole-body control, dexterous manipulation, embodied reasoning, and on-device operation.

The product page presents three related models: Gemini Robotics 2, a vision-language-action model for motor control; Gemini Robotics ER 2, a vision-language model for high-level planning and human communication; and Gemini Robotics On-Device 2, an efficient VLA model that runs locally on robot hardware. Together they are aimed at robots that need to perceive, reason, act, and coordinate across a range of embodiments.

Capabilities

Whole-body robot control

Controls full humanoid robots from feet to fingertips, translating vision and language into motor actions for walking, crouching, reaching, and manipulation.

Dexterous manipulation

Adds a higher level of physical dexterity for both five-finger hands and standard grippers, including delicate and packing-style tasks.

Multi-step embodied reasoning

Lets robots reason about tasks that take minutes and involve many steps, then coordinate execution with the VLA model while tracking progress.

Multi-robot collaboration

Enables multiple robots to communicate and delegate work to complete a shared mission when one robot is not enough.

On-device adaptation

Runs locally on robotic devices and is designed for fast adaptation to new robot embodiments with only a few hours of data.

Use cases

  • Whole-body household or workspace tasks

    A humanoid robot receives a natural-language instruction, walks to a location, bends or crouches as needed, and places an object in a target bin or shelf.

  • Fine dexterity tasks

    A robot uses a five-finger hand or a parallel gripper to perform detailed manipulation such as knot-tying, ziplock sealing, or tight packing.

  • Multi-step task execution

    A robot plans a longer workflow, tracks whether each step succeeded, retries when needed, and keeps going until the task is complete.

  • Collaborative robot teams

    Multiple robots coordinate in the same environment, communicate their strengths, and divide work to finish a shared mission.

  • On-device deployment

    A local robot deployment adapts a model to a new embodiment without relying on network connectivity or cloud latency.

Pros and Cons

Pros

  • Controls full humanoid bodies rather than only upper-body motions.
  • Supports both hands and grippers across different robot embodiments.
  • Combines planning, progress tracking, and action execution in a split ER/VLA workflow.
  • Can run locally on-device for deployments that need low latency or no internet connectivity.
  • Designed to adapt to new robot bodies with a few hours of data.

Cons

  • The source says movement speed still has room to improve.
  • Multi-finger dexterous manipulation is described as remaining challenging.
  • Access is limited to Google AI Studio, private preview, and early-access partners for the models mentioned.

FAQ

What is Gemini Robotics ER 2 used for?

Gemini Robotics ER 2 is the embodied reasoning model. It understands the physical world, plans multi-step tasks, communicates with humans, and hands off motor execution to a lower-level VLA model.

How is Gemini Robotics 2 accessed?

The source says Gemini Robotics 2 is available as a vision-language-action model for controlling robots, Gemini Robotics ER 2 is available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform, and the VLA and On-Device models are available to early-access partners.

How do the ER and VLA models work together?

The pages describe a workflow where ER 2 handles reasoning, planning, progress tracking, and coordination, while the VLA model performs motor control on the robot.

What does the on-device model add?

The on-device model is optimized to run locally on robotic devices and is designed for fast adaptation to new robot embodiments with a few hours of data.

What limitations are mentioned?

The source notes that whole-body control and medium-to-high success rates are demonstrated for whole-body and gripper-based dexterous tasks, while multi-finger dexterous manipulation remains challenging.

Quick Facts

Category
World models & physical AI
Product family
Gemini Robotics 2
Model types
VLA, embodied reasoning (ER), and on-device VLA
Access
Google AI Studio, Gemini Enterprise Agent Platform private preview, and early-access partners
Source domain
deepmind.google