$ cat wiki/models/gemini-robotics-2.md
Gemini Robotics 2
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-07-28 |
| Announced | 2026-07-28 |
| Context window | unknown |
| Pricing | unknown |
| License | unknown |
| Availability | early-access partnership only (not public) |
| Full name | Gemini Robotics 2 (vision-language-action model) |
| Architecture | VLA — single learned policy over the full humanoid |
Release Date
Announced 2026-07-28 as the vision-language-action member of a three-model family. The embodied-reasoning sibling Gemini Robotics ER 2 followed on 2026-07-30 and is the only one of the three available publicly (source).
| Model | Type | Availability |
|---|---|---|
| Gemini Robotics 2 | VLA | early-access partnership |
| Gemini Robotics ER 2 | embodied-reasoning VLM | public (Gemini API, AI Studio) |
| Gemini Robotics On-Device 2 | on-device VLA | early-access partnership |
Benchmarks
DeepMind describes this as its first AI to control a full humanoid — legs, torso, arms and multi-finger hands — under one learned policy (source).
Whole-body manipulation, on the Apollo 2 humanoid:
| Task | Success |
|---|---|
| Picking from a shelf | 76.3% |
| Picking from the floor | 45.7% |
| Fine-motor dexterity, on the SharpaWave hand (22 degrees of freedom): |
| Task | Success |
|---|---|
| Unscrewing a light bulb | 92% |
| Tying a trash bag | 44% |
| Sealing a ziplock bag | 40% |
| Screwing a bulb in | 36% |
| Dustpan task | 32% |
| The spread is the result. Removing a light bulb is near-solved at 92%; putting one back is 36%, | |
| and the deformable-object tasks (trash bag 44%, ziplock 40%) sit near the bottom. Insertion and | |
| deformable manipulation remain unsolved at this release. |
Gemini Robotics On-Device 2 is reported to adapt to a completely new robot embodiment — different shape, sensors and degrees of freedom — in a few hours with fewer than 200 demonstration examples.
Use Cases
- Humanoid whole-body tasks requiring locomotion and manipulation in one policy: walking, crouching and manipulating while reasoning through the task
- Multi-robot collaboration, orchestrated through Gemini Robotics ER 2
- Cross-embodiment transfer via the On-Device 2 variant
Compared To
| Model | Scope | Availability | Developer |
|---|---|---|---|
| Gemini Robotics 2 | full humanoid, one policy | early access | Google DeepMind |
| Gemini Robotics ER 1.6 | embodied reasoning only | Gemini API | Google DeepMind |
| Robostral Navigate | 8B navigation, single RGB camera | — | Mistral AI |
| Cosmos-H-Dreams | surgical robotics world model | Apache 2.0 | NVIDIA |
| No shared benchmark connects this release to Gemini Robotics ER 1.6 — the ER 1.6 | |||
| headline figure (gauge reading 23% → 93%) has no counterpart here, so the two cannot be ranked | |||
| against each other (source). |
Safety
Released alongside ASIMOV-Agentic, an open benchmark for agentic robotic safety — see Gemini Robotics ER 2, where the evaluated behaviours sit.