Google RT-2-X

Robotica / VLANiet beschikbaar
van Google DeepMindModel-ID: rt-2-x

Google's VLA from RT-X collaboration. Trained on Open-X-Embodiment (22 robots, 527 skills), positive transfer.

Status
Niet beschikbaar
Invoer → Uitvoer
Tekst + Afbeelding → Robotacties
Ontwikkelaar
Google DeepMind
Bijgewerkt
23 september 2026

Google RT-2-X is momenteel niet beschikbaar

Je kunt de details op deze pagina nog steeds lezen. Kies een van de beschikbare alternatieven hieronder om direct een vergelijkbaar model uit te voeren.

01

Playground

Google RT-2-X

Onderzoeksmodel

Momenteel niet beschikbaar

Google RT-2-X is een roboticamodel (vision-language-action) en kan niet via de railwail-API worden uitgevoerd.

02

Over Google RT-2-X

SamengevatPer 23 september 2026

Google RT-2-X is een model van Google DeepMind in de categorie Robotica / VLA. Google RT-2-X is momenteel niet beschikbaar op Railwail.

Achtergrond

Over Google DeepMind

Opgericht 2010 · London, UK / Mountain View, USA

RT-2-X is Google DeepMind's flagship Robotic Transformer 2 (RT-2) model retrained on the Open-X-Embodiment dataset - the first large-scale, multi-institution effort to assemble a unified robot-learning dataset spanning many labs and robots. Open-X-Embodiment was organised in 2023 by Google DeepMind together with 21+ academic and industry institutions (Stanford, UC Berkeley, CMU, MIT, Toyota Research Institute, etc.), producing the RT-X dataset of ~1 million trajectories across 22 robot embodiments. RT-2-X extends RT-2's Vision-Language-Action recipe - using a PaLM-E / PaLI-X style VLM as backbone and emitting actions as text tokens - to this cross-embodiment corpus, demonstrating positive transfer across robots and a new state of the art on generalist manipulation at the time of release. RT-2-X is research-only and not publicly callable; it remains a key academic reference and the conceptual parent of subsequent open VLAs.

Google DeepMind bezoeken

Architectuur

Vision-Language-Action transformer (PaLI / PaLM-E backbone, discrete action tokens)

RT-2-X follows the RT-2 design: a large Vision-Language Model (PaLI-X or PaLM-E) is co-fine-tuned on web-scale vision-language data and on robot demonstration data, where robot actions are tokenised as strings of natural-language-like tokens (each action dimension binned and rendered as a token). The same next-token prediction objective therefore trains the model on both internet-scale image-text data and on robot trajectories, allowing the resulting policy to inherit web knowledge (object semantics, OCR, common sense) and route it to motor commands. RT-2-X is the version of this recipe trained on the Open-X-Embodiment / RT-X dataset - ~1 million trajectories across 22 robot embodiments - rather than only on Google's internal kitchen-robot dataset. Public results report 5B and 55B variants, with the 55B model showing the strongest generalisation, especially when prompted with unseen language commands or unseen object combinations.

Parameters
Up to 55B (RT-2-X variants: 5B and 55B)

Mogelijkheden

  • Generalist VLA trained on Open-X-Embodiment (22 robots)
  • Inherits web-scale knowledge from PaLI / PaLM-E backbones
  • Discrete action-token decoding (text-like vocabulary)
  • Positive transfer across robot embodiments
  • Strong emergent semantic reasoning (e.g. 'pick up the extinct animal')
  • 5B and 55B parameter variants
  • Reference architecture for the modern VLA paradigm
  • Co-training on internet data + robot demos
  • Best for: research, citation, conceptual baseline for VLAs.

Training & licentie

Co-trained on internet-scale vision-language data (PaLI / PaLM-E corpora) plus ~1 million robot trajectories from the Open-X-Embodiment (RT-X) dataset across 22 robot embodiments. Action targets are tokenised continuous controls.

Licentie: Research-only - Google DeepMind has not publicly released the RT-2-X weights, code or API. Some Open-X-Embodiment data and smaller RT-X reproductions are available, but the proprietary RT-2-X checkpoints are not.

Veiligheidstests: Covered under Google's standard Responsible AI process and frontier-safety evaluations. Robot demos were filmed in supervised lab settings with force-limited hardware; no public RSP-style document specific to RT-2-X.

Bekende beperkingen

  • Closed weights - no public API or download
  • Inference latency too high for very fast control loops
  • Discrete tokens limit smoothness vs diffusion / flow-matching policies
  • Cross-embodiment transfer still constrained by action-space differences
  • Long-horizon tasks need external prompt decomposition
  • Dataset skew toward kitchen / tabletop tasks
03

Prijzen

Momenteel niet beschikbaar. Er is momenteel geen prijs voor dit model, dus het kan niet worden uitgevoerd.

04

API

Roep Google RT-2-X aan met je Railwail API-sleutel. Gebruik deze model-ID in het verzoek:

Niet beschikbaar via de API

Robotica-modellen draaien op robotica-hardware, niet via de railwail-API.

05

Specificaties

Model-ID
rt-2-x
Ontwikkelaar
Google DeepMind
Invoer
Tekst, Afbeelding
Uitvoer
Robotacties
Levenscyclus
Niet beschikbaar
Modelgrootte
Up to 55B (RT-2-X variants: 5B and 55B)
Licentie
Research-only - Google DeepMind has not publicly released the RT-2-X weights, code or API. Some Open-X-Embodiment data and smaller RT-X reproductions are available, but the proprietary RT-2-X checkpoints are not.
Catalogusitem bijgewerkt
23 september 2026

Tags

  • google
  • vla
  • robotics
  • research-only
  • weights-closed
06

Gebruiksscenario's

Waarvoor het wordt gebruikt

  • Reference baseline in VLA / cross-embodiment papers
  • Comparative studies vs OpenVLA / Octo / Ï€-0
  • Academic citation as the canonical large VLA
  • Demonstration of emergent reasoning in robotics
  • Motivating examples for action-token tokenisation research
  • Closed model - no direct deployment use
07

Veelgestelde vragen

Wat is Google RT-2-X?

Google RT-2-X is een model van Google DeepMind in de categorie Robotica / VLA. Het staat in de Railwail-catalogus, maar kan momenteel niet worden uitgevoerd.

Hoeveel kost Google RT-2-X op Railwail?

Google RT-2-X kan momenteel niet op Railwail worden uitgevoerd, dus er is geen huidige prijs. Beschikbare alternatieven met prijzen staan verderop op deze pagina.

Hoe snel is Google RT-2-X?

Er zijn nog niet genoeg gemeten runs van Google RT-2-X op Railwail om een uitvoeringstijd op te geven. Dit hangt af van de invoer, de instellingen en de belasting bij de provider.

Wanneer moet ik Google RT-2-X gebruiken?

Google RT-2-X behoort tot de categorie Robotica / VLA. De categoriepagina bevat de andere modellen van dit type met hun prijzen.

Alle modellen: Robotica / VLA

Kan Google RT-2-X afbeeldingen verwerken?

Ja. Google RT-2-X accepteert afbeeldingen als invoer naast tekst.

Kan ik Google RT-2-X nu gebruiken?

Momenteel niet beschikbaar. De pagina blijft online; beschikbare alternatieven uit dezelfde categorie staan verderop.

Alle modellen via één API

Één API-sleutel voor elk model op Railwail. Gebruik wordt afgerekend via vooraf gekochte credits, 1 credit = US$ 0,01.