Google DeepMind has released Gemini Robotics 2, a family of three robot AI models that splits the work between a model that understands a scene and plans, and models that turn those plans into motion from fingertips to feet. Only the planning model is open to the public through an API; the motion models are limited to trusted testers. The research lead for Gemini Robotics at Google DeepMind says the pace of progress is surprising, yet still places robotics at roughly the “GPT-2 era” of language models. Here is what changed, what still holds robots back, and what teams evaluating robot AI can do now.

Three models, two layers

Gemini Robotics 2 is built as a stack. A slower reasoning layer decides what to do, and an action layer decides how the body moves.

Model What it does Availability
Gemini Robotics ER 2 Scene understanding, spatial analysis, task planning, tool selection Public API in Google AI Studio
Gemini Robotics 2 Vision-language-action model that turns instructions into joint motions and trajectories Trusted testers only
Gemini Robotics On-Device 2 Compact action model that runs on the robot without a cloud connection Trusted testers only

ER stands for embodied reasoning: reasoning about a physical scene from the point of view of a body that has to act in it. Gemini Robotics ER 2 is derived from Gemini Flash and works as the “System 2” planner, the slow and deliberate thinker. Developers describe what their robot can do, such as grabbing or placing an object, as tools with text descriptions. The model then picks the right tool, much as a software AI agent picks a function to call.

Gemini Robotics 2 is a VLA, or vision-language-action model: it takes images and language as input and outputs robot actions. It reasons over the whole body at once, coordinating hands, arms, legs and posture rather than treating each joint in isolation. Fast balance and stabilization loops still run in lower-level controllers, but the targets and trajectories come from the VLA. The on-device version is a compressed variant that runs on the robot itself, which removes network round trips and keeps the robot working where bandwidth is poor.

How the planner steers the body

The link between the planner and the action model is a research area called steerability. Earlier versions of ER designated pick-and-place targets with 2D image coordinates. Now the planner can steer the action model in several ways:

  • Natural language, such as “turn around” or “look to your left”
  • Pointing at an object in the scene, which settles ambiguity such as a lemon versus a banana
  • In-context learning, where a demonstration clip shows the robot what to do without any retraining

Passing only plain text between the two models risks losing spatial detail, so the goal is a rich, multimodal interface between them.

Memory is another constraint. The model’s context window, the amount of information it can consider at once, is 128,000 tokens. With video input, that works out to roughly three minutes of episodic memory, depending on frame rate and how images are encoded. That covers short manipulation tasks but not multi-hour chores. A practical answer is to keep compact text summaries of finished steps (“flipped the egg”) instead of raw video, and to downsample frames or drop static ones.

A robot camera view of a lemon and a banana linked to a humanoid body with joints highlighted from fingertips to feet

▲ Planner steering the action model

Why hands are harder than running

Humanoid robots outran top human sprinters at a robot competition in China this summer, and the clips spread widely. From a research point of view, flat-ground running is the easier problem. Contact between feet and a rigid floor is predictable, so simulation models it accurately and skills learned in simulation transfer well to real hardware. The sprinting robots most likely rely on low-level reinforcement learning controllers initialized from human or animal motion data, with little high-level reasoning involved. Speed is also not what makes a robot useful at home or at work.

Manipulation is different. Friction changes, objects bend and fold, and simulators struggle to keep up; folding cloth is a classic example. Rigid pick-and-place is the exception, and foundation models, large models pretrained on broad data, have made it nearly reliable enough for commercial use. That has moved the research frontier from fixed tabletop gripper arms to humanoids that step, squat and handle unfamiliar objects in a closed loop.

Hardware has changed too. Gemini Robotics 1 showed what two-finger parallel grippers could do. About 15 months later, Gemini Robotics 2 showed multi-fingered hands tying a trash bag. Multi-fingered hands now set the frontier for dexterity, and the commercial options vary widely:

Robot hand Size Strength
Wuji Close to a human hand About that of a 10-year-old child
Sharpa Larger than a human hand Lifts about 20 kilograms and can open tight jar lids

The open problems are reliability, consistency from unit to unit, and cost. Touch matters as well. Force sensing and back-drivability, meaning a joint gives way smoothly when pushed, let a robot stack potato chips without crushing them, and feeding fingertip force data into the learning model improves its decisions compared with vision alone.

Why robotics is still in its “GPT-2 era”

GPT-3 marked the point where language models could reliably learn a new task from a few examples. Robotics has not reached that point across a broad range of tasks.

The biggest gap is cross-embodiment: one model controlling robots with different bodies and sensors. A language model runs the same on an iPhone, a Mac or a Linux server, but a robot policy often fails completely on different hardware. Language AI never had to solve this. Google DeepMind is aiming for general intelligence that is abstracted away from any one body, and it uses humanoids as the proving ground because they force progress on three fronts at once: whole-body balance, multi-finger dexterity and human-robot interaction.

There is progress. Google DeepMind works with robot makers including Apptronik, Agile Robots and Boston Dynamics, and the on-device model adapts to a new robot body with about 200 demonstrations. Zero-shot deployment, meaning running on an unfamiliar robot with no extra training and production-grade reliability, is still unsolved. Demonstration-based prompting cuts setup time, but it can slide into rigid imitation when the scene changes.

Errors also compound. In a chain of ten actions, the system has to check whether each step succeeded and choose the next sub-goal. Without solid real-time recovery, overall success drops sharply. Whether a mistake can be undone makes a large difference:

  • LEGO assembly is forgiving: a misplaced brick can simply be tried again.
  • Cooking an egg is not: a dropped or overcooked egg cannot be fixed, so the task demands far higher reliability.

Visual generalization has improved a lot, so changes in lighting or kitchen layout no longer derail models. The bottleneck now is the physical side of tasks: contact, force control and irreversible steps.

A robot arm retrying toy brick assembly beside a robot hand hovering over an egg at the counter edge

▲ Recoverable versus irreversible tasks

Speed and reliability depend on the job

Latency is a real trade-off. In a simple 2D simulation wired to the Gemini Robotics API, the model took about six seconds to go from an image and the instruction “move the banana to the plate” to a tool call. That is slow for responsive interaction, but the acceptable speed depends on the setting:

  • Factory assembly lines cannot pause, so they need very fast responses.
  • Household chores such as folding laundry reward accuracy over speed, and a robot could work slowly overnight.
  • Remote field sites with poor connectivity need models that run on the robot.

Homes still are not the first market. Toddlers and cluttered, unpredictable layouts set a very high safety bar, so controlled commercial settings remain easier for early deployment. Because new capabilities tend to appear first in larger, slower models, fast and slow models are likely to be developed side by side, with knowledge distilled from big models into small ones.

The planner’s advantages are specific. In a similar 2D test of basic pick-and-place, including objects that moved mid-task, Gemini Robotics ER 2 and the base Gemini 1.5 Flash performed equally well. ER 2 pulls ahead on specialized work such as reading analog gauges, inspecting for defects and pointing precisely at locations. After the API went live in Google AI Studio, usage far exceeded what small trusted-tester groups had shown, and researchers mounted it on Boston Dynamics’ Spot to read instruments and hand out snacks. One catch: robotics has no standard APIs, so robots with unusual command sets need repeated prompt and schema tuning before the planner drives them reliably.

Safety as a capability

A chatbot that is right half the time can still be useful. A robot that drops things causes real damage, so success rates of 50% or even 80% are not enough for autonomous work unless the task tolerates failure. Safety is treated as a capability in its own right, not a brake on capability, because consumers will not accept home robots that are not safe and reliable.

Robot safety has two parts: physical safety, which covers clumsiness, mechanical failure and tipping over, and behavioral alignment of the models that plan. A heavy humanoid falling inside a home is an immediate hazard, so safety has to run through the whole stack, from mechanical design and emergency-stop hardware up to the planning model.

Emergency stop How it works Effect on a legged robot
Hard e-stop Cuts power to every joint The body can collapse onto nearby people or objects
Soft e-stop Freezes joints in place The robot holds its balance and does not fall

Behavior gets tested too. In one robustness test, a researcher put a basket over a working humanoid’s head. The right response is to notice the blocked sensors, stop and ask a person for help rather than keep moving blind. The human shape also raises the bar: when a gripper misses an object, people see a glitch, but when a humanoid fumbles, it looks far less competent. As robots gain natural behaviors, such as gestures generated on the fly in conversation instead of scripted nods, those expectations may climb further.

What to take away

Gemini Robotics 2 separates planning from motion and opens the planner to anyone through an API, while dexterous manipulation, cross-embodiment, compounding errors and safety remain open research problems. For developers and companies evaluating robot AI, these steps follow from where the technology stands:

  1. Start with what is available: the Gemini Robotics ER 2 API. Define robot actions as clearly named and described tools, let ER 2 plan, and leave motion to existing controllers or dedicated action models.
  2. Compare ER 2 with a general model on the tasks where it is strongest, such as gauge reading, inspection and pointing, not just basic pick-and-place.
  3. Set latency requirements first. A factory line and an overnight household chore call for different models and deployment choices.
  4. Begin with tasks that can be retried after a failure, and add a success check between steps in any multi-step sequence.
  5. Test with moved objects and changed scenes, not the training setup, to tell real generalization from imitation.
  6. For legged robots, use stops that lock joints rather than cut power, and make the robot halt and ask for help when its sensors are blocked.