Gemini Robotics-ER

Robotik / VLAInte tillgÀnglig
av Google DeepMindModell-ID: gemini-robotics-er

Embodied-reasoning variant of Gemini Robotics. Enhanced 3D spatial reasoning and trajectory planning.

Status
Inte tillgÀnglig
Inmatning → utmatning
Text + Bild → Roboteraktioner
Utvecklare
Google DeepMind
Uppdaterad
23 september 2026

Gemini Robotics-ER Àr för nÀrvarande otillgÀnglig

Du kan fortfarande lÀsa detaljerna pÄ denna sida. VÀlj ett av de tillgÀngliga alternativen nedan för att köra en jÀmförbar modell direkt.

01

Playground

Gemini Robotics-ER

Forskningsmodell

Inte tillgÀnglig för nÀrvarande

Gemini Robotics-ER Àr en robotikmodell (vision-language-action) och kan inte köras via railwail API.

02

Om Gemini Robotics-ER

Kort sagtFrÄn och med 23 september 2026

Gemini Robotics-ER Àr en modell av Google DeepMind i kategorin Robotik / VLA. Gemini Robotics-ER Àr för nÀrvarande inte tillgÀnglig pÄ Railwail.

Bakgrund

Om Google DeepMind

Grundat 2010 · London, UK / Mountain View, USA

Google DeepMind announced Gemini Robotics-ER (Embodied Reasoning) alongside Gemini Robotics in March 2025. While Gemini Robotics is the action-producing VLA, Gemini Robotics-ER is the reasoning-focused sibling: a Vision-Language Model variant of Gemini 2.0 specialised for spatial understanding, 3D grounding, point/box prediction, trajectory planning and code generation for robotics. It is designed to be combined with classical motion planners, low-level controllers or with the Gemini Robotics VLA itself. DeepMind positions Gemini Robotics-ER as a 'reasoning brain' that a robot stack can call with multimodal prompts to decompose tasks, locate objects in 2D / 3D, and emit waypoints or Python control code. As with Gemini Robotics, access is limited to research and partner programs.

Besök Google DeepMind

Arkitektur

Vision-Language Model for Embodied Reasoning (no end-to-end action head)

Gemini Robotics-ER is a fine-tuned variant of Gemini 2.0 specialised for embodied perception and planning rather than direct control. The architecture preserves the multimodal Transformer backbone of Gemini 2.0 (image, video, text, code) but is post-trained on a curated corpus of embodied tasks: object detection in 2D and 3D, point and bounding-box prediction, grasp prediction, motion-trajectory generation, and code-as-policy outputs that call robot APIs. It can accept egocentric robot camera streams and a natural-language task description, then produce structured outputs such as pixel-space points to grasp, 3D coordinates relative to the camera, planning steps, or Python snippets that drive a downstream controller. In combination with the Gemini Robotics VLA, Robotics-ER provides high-level reasoning while the VLA handles closed-loop low-level actions.

Parameter
Undisclosed (Gemini 2.0-class)

Funktioner

  • Spatial reasoning over 2D / 3D scenes
  • Point and bounding-box prediction for objects and grasps
  • Trajectory waypoint generation
  • Code-as-policy generation (Python that calls robot APIs)
  • Compositional task planning from natural language
  • Pair with motion planners or with Gemini Robotics VLA
  • Multimodal context: images, video, text, robot state
  • Improved zero-shot performance on embodied QA benchmarks
  • Best for: planning and grounding modules in research robotics stacks.

TrÀning & licens

Gemini 2.0 multimodal pretraining plus embodied post-training on object-detection, 3D grounding, grasp prediction, trajectory planning, and code-generation tasks for robotic control. Draws on Google's internal robot datasets and curated public embodied datasets.

Licens: Research-only / partner access through Google DeepMind. Not publicly downloadable.

SĂ€kerhetstestning: Covered under Google DeepMind's Frontier Safety Framework and Responsible AI evaluations, with specific embodied-AI red-teaming for physical-harm and unsafe-action scenarios.

KÀnda begrÀnsningar

  • No direct low-level action output
  • Requires downstream controller or planner
  • Closed model - no public weights or API
  • Spatial reasoning still imperfect on cluttered scenes
  • Latency too high for tight inner control loops
  • Generalisation depends on prompt and tool stack
03

Priser

Inte tillgÀnglig för nÀrvarande. Det finns ingen pris för denna modell för nÀrvarande, sÄ den kan inte köras.

04

API

Anropa Gemini Robotics-ER med din Railwail API-nyckel. AnvÀnd detta modell-ID i begÀran:

Inte tillgÀnglig via API:et

Robotikmodeller körs pÄ robotmaskinvara, inte via railwail-API:et.

05

Specifikationer

Modell-ID
gemini-robotics-er
Utvecklare
Google DeepMind
Inmatning
Text, Bild
Utmatning
Roboteraktioner
Livscykel
Aktuell version
Modellstorlek
Undisclosed (Gemini 2.0-class)
Licens
Research-only / partner access through Google DeepMind. Not publicly downloadable.
KataloginlÀgg uppdaterat
23 september 2026

Taggar

  • google
  • deepmind
  • gemini
  • vla
  • robotics
  • research-only
  • weights-closed
  • embodied-reasoning
06

AnvÀndningsfall

Vad det anvÀnds till

  • High-level planner in hierarchical robot stacks
  • Grounding language to 2D / 3D scene primitives
  • Code-as-policy generation for manipulation
  • Embodied QA and instruction parsing
  • Companion to motion planners or VLA controllers
  • Academic embodied-reasoning research
07

Vanliga frÄgor

Vad Àr Gemini Robotics-ER?

Gemini Robotics-ER Àr en modell av Google DeepMind i kategorin Robotik / VLA. Den finns i Railwail-katalogen men kan inte köras för nÀrvarande.

Vad kostar Gemini Robotics-ER pÄ Railwail?

Gemini Robotics-ER kan inte köras pÄ Railwail för nÀrvarande, sÄ det finns inget aktuellt pris. TillgÀngliga alternativ med priser visas lÀngre ned pÄ denna sida.

Hur snabb Àr Gemini Robotics-ER?

Det finns Ànnu inte tillrÀckligt mÄnga uppmÀtta körningar av Gemini Robotics-ER pÄ Railwail för att ange en körningstid. Det beror pÄ inmatningen, instÀllningarna och belastningen hos leverantören.

NÀr ska jag anvÀnda Gemini Robotics-ER?

Gemini Robotics-ER tillhör kategorin Robotik / VLA. Kategorisidan listar de andra modellerna av denna typ med deras priser.

Alla modeller: Robotik / VLA

Kan Gemini Robotics-ER bearbeta bilder?

Ja. Gemini Robotics-ER accepterar bilder som inmatning utöver text.

Kan jag anvÀnda Gemini Robotics-ER just nu?

Inte tillgÀnglig för nÀrvarande. Sidan förblir online; tillgÀngliga alternativ frÄn samma kategori visas lÀngre ned.

Alla modeller via ett API

En API-nyckel för alla modeller pÄ Railwail. AnvÀndningen debiteras frÄn förbetald kredit, 1 kredit = 0,01 US$.