Gemini Robotics-ER

Robotik / VLANicht verfügbar
von Google DeepMindModell-ID: gemini-robotics-er

Embodied-reasoning variant of Gemini Robotics. Enhanced 3D spatial reasoning and trajectory planning.

Status
Nicht verfügbar
Eingabe → Ausgabe
Text + Bild → Roboteraktionen
Entwickler
Google DeepMind
Aktualisiert
23. September 2026

Gemini Robotics-ER ist derzeit nicht verfügbar

Die Angaben auf dieser Seite kannst du weiter nachlesen. Mit einer der verfügbaren Alternativen unten kannst du sofort ein vergleichbares Modell nutzen.

01

Playground

Gemini Robotics-ER

Forschungsmodell

Derzeit nicht verfügbar

Gemini Robotics-ER ist ein Robotik-Modell (Vision-Language-Action) und lässt sich nicht über die railwail-API ausführen.

02

Über Gemini Robotics-ER

Kurz gesagtStand: 23. September 2026

Gemini Robotics-ER ist ein Modell von Google DeepMind aus der Kategorie Robotik / VLA. Über Railwail ist Gemini Robotics-ER derzeit nicht verfügbar.

Hintergrund

Über Google DeepMind

Gegründet 2010 · London, UK / Mountain View, USA

Google DeepMind announced Gemini Robotics-ER (Embodied Reasoning) alongside Gemini Robotics in March 2025. While Gemini Robotics is the action-producing VLA, Gemini Robotics-ER is the reasoning-focused sibling: a Vision-Language Model variant of Gemini 2.0 specialised for spatial understanding, 3D grounding, point/box prediction, trajectory planning and code generation for robotics. It is designed to be combined with classical motion planners, low-level controllers or with the Gemini Robotics VLA itself. DeepMind positions Gemini Robotics-ER as a 'reasoning brain' that a robot stack can call with multimodal prompts to decompose tasks, locate objects in 2D / 3D, and emit waypoints or Python control code. As with Gemini Robotics, access is limited to research and partner programs.

Google DeepMind besuchen

Architektur

Vision-Language Model for Embodied Reasoning (no end-to-end action head)

Gemini Robotics-ER is a fine-tuned variant of Gemini 2.0 specialised for embodied perception and planning rather than direct control. The architecture preserves the multimodal Transformer backbone of Gemini 2.0 (image, video, text, code) but is post-trained on a curated corpus of embodied tasks: object detection in 2D and 3D, point and bounding-box prediction, grasp prediction, motion-trajectory generation, and code-as-policy outputs that call robot APIs. It can accept egocentric robot camera streams and a natural-language task description, then produce structured outputs such as pixel-space points to grasp, 3D coordinates relative to the camera, planning steps, or Python snippets that drive a downstream controller. In combination with the Gemini Robotics VLA, Robotics-ER provides high-level reasoning while the VLA handles closed-loop low-level actions.

Parameter
Undisclosed (Gemini 2.0-class)

Fähigkeiten

  • Spatial reasoning over 2D / 3D scenes
  • Point and bounding-box prediction for objects and grasps
  • Trajectory waypoint generation
  • Code-as-policy generation (Python that calls robot APIs)
  • Compositional task planning from natural language
  • Pair with motion planners or with Gemini Robotics VLA
  • Multimodal context: images, video, text, robot state
  • Improved zero-shot performance on embodied QA benchmarks
  • Best for: planning and grounding modules in research robotics stacks.

Training & Lizenz

Gemini 2.0 multimodal pretraining plus embodied post-training on object-detection, 3D grounding, grasp prediction, trajectory planning, and code-generation tasks for robotic control. Draws on Google's internal robot datasets and curated public embodied datasets.

Lizenz: Research-only / partner access through Google DeepMind. Not publicly downloadable.

Sicherheitstests: Covered under Google DeepMind's Frontier Safety Framework and Responsible AI evaluations, with specific embodied-AI red-teaming for physical-harm and unsafe-action scenarios.

Bekannte Grenzen

  • No direct low-level action output
  • Requires downstream controller or planner
  • Closed model - no public weights or API
  • Spatial reasoning still imperfect on cluttered scenes
  • Latency too high for tight inner control loops
  • Generalisation depends on prompt and tool stack
03

Preise

Derzeit nicht verfügbar. Für dieses Modell gibt es derzeit keinen Preis, deshalb lässt es sich nicht ausführen.

04

API

Rufe Gemini Robotics-ER mit deinem Railwail-API-Schlüssel auf. Diese Modell-ID gehört in die Anfrage:

Nicht über die API verfügbar

Robotik-Modelle laufen auf Roboter-Hardware, nicht über die railwail-API.

05

Spezifikationen

Modell-ID
gemini-robotics-er
Entwickler
Google DeepMind
Kategorie
Robotik / VLA
Eingabe
Text, Bild
Ausgabe
Roboteraktionen
Lebenszyklus
Aktuelle Version
Modellgröße
Undisclosed (Gemini 2.0-class)
Lizenz
Research-only / partner access through Google DeepMind. Not publicly downloadable.
Katalogeintrag aktualisiert
23. September 2026

Schlagwörter

  • google
  • deepmind
  • gemini
  • vla
  • robotics
  • research-only
  • weights-closed
  • embodied-reasoning
06

Einsatzgebiete

Wofür es genutzt wird

  • High-level planner in hierarchical robot stacks
  • Grounding language to 2D / 3D scene primitives
  • Code-as-policy generation for manipulation
  • Embodied QA and instruction parsing
  • Companion to motion planners or VLA controllers
  • Academic embodied-reasoning research
07

Häufige Fragen

Was ist Gemini Robotics-ER?

Gemini Robotics-ER ist ein Modell von Google DeepMind aus der Kategorie Robotik / VLA. Es steht im Railwail-Katalog, lässt sich derzeit aber nicht ausführen.

Was kostet Gemini Robotics-ER bei Railwail?

Gemini Robotics-ER lässt sich über Railwail derzeit nicht ausführen, deshalb gibt es keinen aktuellen Preis. Verfügbare Alternativen mit Preisen stehen weiter unten auf dieser Seite.

Wie schnell ist Gemini Robotics-ER?

Für Gemini Robotics-ER gibt es bei Railwail noch zu wenige gemessene Läufe, um eine Laufzeit anzugeben. Sie hängt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Wann ist Gemini Robotics-ER die richtige Wahl?

Gemini Robotics-ER gehört zur Kategorie Robotik / VLA. Die Kategorieseite listet die anderen Modelle dieser Art mit ihren Preisen.

Alle Modelle: Robotik / VLA

Kann Gemini Robotics-ER Bilder verarbeiten?

Ja. Gemini Robotics-ER nimmt neben Text auch Bilder als Eingabe an.

Kann ich Gemini Robotics-ER gerade nutzen?

Derzeit nicht verfügbar. Die Seite bleibt online; verfügbare Alternativen aus derselben Kategorie stehen weiter unten.

Alle Modelle über eine API

Ein API-Schlüssel für alle Modelle auf Railwail. Abgerechnet wird über vorab gekaufte Credits, 1 Credit = 0,01 $.