Gemini Robotics-ER

Robotics / VLAUnavailable
by Google DeepMindModel ID: gemini-robotics-er

Embodied-reasoning variant of Gemini Robotics. Enhanced 3D spatial reasoning and trajectory planning.

Status
Unavailable
Input โ†’ output
Text + Image โ†’ Robot actions
Developer
Google DeepMind
Updated
September 23, 2026

Gemini Robotics-ER is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

01

Playground

Gemini Robotics-ER

Research model

Currently unavailable

Gemini Robotics-ER is a robotics model (vision-language-action) and cannot be run through the railwail API.

02

About Gemini Robotics-ER

TL;DRAs of September 23, 2026

Gemini Robotics-ER is a model by Google DeepMind in the Robotics / VLA category. Gemini Robotics-ER is currently not available on Railwail.

Background

About Google DeepMind

Founded 2010 ยท London, UK / Mountain View, USA

Google DeepMind announced Gemini Robotics-ER (Embodied Reasoning) alongside Gemini Robotics in March 2025. While Gemini Robotics is the action-producing VLA, Gemini Robotics-ER is the reasoning-focused sibling: a Vision-Language Model variant of Gemini 2.0 specialised for spatial understanding, 3D grounding, point/box prediction, trajectory planning and code generation for robotics. It is designed to be combined with classical motion planners, low-level controllers or with the Gemini Robotics VLA itself. DeepMind positions Gemini Robotics-ER as a 'reasoning brain' that a robot stack can call with multimodal prompts to decompose tasks, locate objects in 2D / 3D, and emit waypoints or Python control code. As with Gemini Robotics, access is limited to research and partner programs.

Visit Google DeepMind

Architecture

Vision-Language Model for Embodied Reasoning (no end-to-end action head)

Gemini Robotics-ER is a fine-tuned variant of Gemini 2.0 specialised for embodied perception and planning rather than direct control. The architecture preserves the multimodal Transformer backbone of Gemini 2.0 (image, video, text, code) but is post-trained on a curated corpus of embodied tasks: object detection in 2D and 3D, point and bounding-box prediction, grasp prediction, motion-trajectory generation, and code-as-policy outputs that call robot APIs. It can accept egocentric robot camera streams and a natural-language task description, then produce structured outputs such as pixel-space points to grasp, 3D coordinates relative to the camera, planning steps, or Python snippets that drive a downstream controller. In combination with the Gemini Robotics VLA, Robotics-ER provides high-level reasoning while the VLA handles closed-loop low-level actions.

Parameters
Undisclosed (Gemini 2.0-class)

Capabilities

  • Spatial reasoning over 2D / 3D scenes
  • Point and bounding-box prediction for objects and grasps
  • Trajectory waypoint generation
  • Code-as-policy generation (Python that calls robot APIs)
  • Compositional task planning from natural language
  • Pair with motion planners or with Gemini Robotics VLA
  • Multimodal context: images, video, text, robot state
  • Improved zero-shot performance on embodied QA benchmarks
  • Best for: planning and grounding modules in research robotics stacks.

Training & license

Gemini 2.0 multimodal pretraining plus embodied post-training on object-detection, 3D grounding, grasp prediction, trajectory planning, and code-generation tasks for robotic control. Draws on Google's internal robot datasets and curated public embodied datasets.

License: Research-only / partner access through Google DeepMind. Not publicly downloadable.

Safety testing: Covered under Google DeepMind's Frontier Safety Framework and Responsible AI evaluations, with specific embodied-AI red-teaming for physical-harm and unsafe-action scenarios.

Known limitations

  • No direct low-level action output
  • Requires downstream controller or planner
  • Closed model - no public weights or API
  • Spatial reasoning still imperfect on cluttered scenes
  • Latency too high for tight inner control loops
  • Generalisation depends on prompt and tool stack
03

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

04

API

Call Gemini Robotics-ER with your Railwail API key. Use this model ID in the request:

Not available via the API

Robotics models run on robot hardware, not through the railwail API.

05

Specifications

Model ID
gemini-robotics-er
Input
Text, Image
Output
Robot actions
Lifecycle
Current version
Model size
Undisclosed (Gemini 2.0-class)
License
Research-only / partner access through Google DeepMind. Not publicly downloadable.
Catalog entry updated
September 23, 2026

Tags

  • google
  • deepmind
  • gemini
  • vla
  • robotics
  • research-only
  • weights-closed
  • embodied-reasoning
06

Use cases

What it is used for

  • High-level planner in hierarchical robot stacks
  • Grounding language to 2D / 3D scene primitives
  • Code-as-policy generation for manipulation
  • Embodied QA and instruction parsing
  • Companion to motion planners or VLA controllers
  • Academic embodied-reasoning research
07

Frequently asked questions

What is Gemini Robotics-ER?

Gemini Robotics-ER is a model by Google DeepMind in the Robotics / VLA category. It is listed on Railwail but cannot be run at the moment.

How much does Gemini Robotics-ER cost on Railwail?

Gemini Robotics-ER cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

How fast is Gemini Robotics-ER?

There are not enough measured runs of Gemini Robotics-ER on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

When should I use Gemini Robotics-ER?

Gemini Robotics-ER belongs to the Robotics / VLA category. The category page lists the other models of this kind with their prices.

All models in Robotics / VLA

Can Gemini Robotics-ER process images?

Yes. Gemini Robotics-ER accepts images as input in addition to text.

Can I use Gemini Robotics-ER right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.