Google RT-2-X

Robotik / VLAInte tillgÀnglig
av Google DeepMindModell-ID: rt-2-x

Google's VLA from RT-X collaboration. Trained on Open-X-Embodiment (22 robots, 527 skills), positive transfer.

Status
Inte tillgÀnglig
Inmatning → utmatning
Text + Bild → Roboteraktioner
Utvecklare
Google DeepMind
Uppdaterad
23 september 2026

Google RT-2-X Àr för nÀrvarande otillgÀnglig

Du kan fortfarande lÀsa detaljerna pÄ denna sida. VÀlj ett av de tillgÀngliga alternativen nedan för att köra en jÀmförbar modell direkt.

01

Playground

Google RT-2-X

Forskningsmodell

Inte tillgÀnglig för nÀrvarande

Google RT-2-X Àr en robotikmodell (vision-language-action) och kan inte köras via railwail API.

02

Om Google RT-2-X

Kort sagtFrÄn och med 23 september 2026

Google RT-2-X Àr en modell av Google DeepMind i kategorin Robotik / VLA. Google RT-2-X Àr för nÀrvarande inte tillgÀnglig pÄ Railwail.

Bakgrund

Om Google DeepMind

Grundat 2010 · London, UK / Mountain View, USA

RT-2-X is Google DeepMind's flagship Robotic Transformer 2 (RT-2) model retrained on the Open-X-Embodiment dataset - the first large-scale, multi-institution effort to assemble a unified robot-learning dataset spanning many labs and robots. Open-X-Embodiment was organised in 2023 by Google DeepMind together with 21+ academic and industry institutions (Stanford, UC Berkeley, CMU, MIT, Toyota Research Institute, etc.), producing the RT-X dataset of ~1 million trajectories across 22 robot embodiments. RT-2-X extends RT-2's Vision-Language-Action recipe - using a PaLM-E / PaLI-X style VLM as backbone and emitting actions as text tokens - to this cross-embodiment corpus, demonstrating positive transfer across robots and a new state of the art on generalist manipulation at the time of release. RT-2-X is research-only and not publicly callable; it remains a key academic reference and the conceptual parent of subsequent open VLAs.

Besök Google DeepMind

Arkitektur

Vision-Language-Action transformer (PaLI / PaLM-E backbone, discrete action tokens)

RT-2-X follows the RT-2 design: a large Vision-Language Model (PaLI-X or PaLM-E) is co-fine-tuned on web-scale vision-language data and on robot demonstration data, where robot actions are tokenised as strings of natural-language-like tokens (each action dimension binned and rendered as a token). The same next-token prediction objective therefore trains the model on both internet-scale image-text data and on robot trajectories, allowing the resulting policy to inherit web knowledge (object semantics, OCR, common sense) and route it to motor commands. RT-2-X is the version of this recipe trained on the Open-X-Embodiment / RT-X dataset - ~1 million trajectories across 22 robot embodiments - rather than only on Google's internal kitchen-robot dataset. Public results report 5B and 55B variants, with the 55B model showing the strongest generalisation, especially when prompted with unseen language commands or unseen object combinations.

Parameter
Up to 55B (RT-2-X variants: 5B and 55B)

Funktioner

  • Generalist VLA trained on Open-X-Embodiment (22 robots)
  • Inherits web-scale knowledge from PaLI / PaLM-E backbones
  • Discrete action-token decoding (text-like vocabulary)
  • Positive transfer across robot embodiments
  • Strong emergent semantic reasoning (e.g. 'pick up the extinct animal')
  • 5B and 55B parameter variants
  • Reference architecture for the modern VLA paradigm
  • Co-training on internet data + robot demos
  • Best for: research, citation, conceptual baseline for VLAs.

TrÀning & licens

Co-trained on internet-scale vision-language data (PaLI / PaLM-E corpora) plus ~1 million robot trajectories from the Open-X-Embodiment (RT-X) dataset across 22 robot embodiments. Action targets are tokenised continuous controls.

Licens: Research-only - Google DeepMind has not publicly released the RT-2-X weights, code or API. Some Open-X-Embodiment data and smaller RT-X reproductions are available, but the proprietary RT-2-X checkpoints are not.

SĂ€kerhetstestning: Covered under Google's standard Responsible AI process and frontier-safety evaluations. Robot demos were filmed in supervised lab settings with force-limited hardware; no public RSP-style document specific to RT-2-X.

KÀnda begrÀnsningar

  • Closed weights - no public API or download
  • Inference latency too high for very fast control loops
  • Discrete tokens limit smoothness vs diffusion / flow-matching policies
  • Cross-embodiment transfer still constrained by action-space differences
  • Long-horizon tasks need external prompt decomposition
  • Dataset skew toward kitchen / tabletop tasks
03

Priser

Inte tillgÀnglig för nÀrvarande. Det finns ingen pris för denna modell för nÀrvarande, sÄ den kan inte köras.

04

API

Anropa Google RT-2-X med din Railwail API-nyckel. AnvÀnd detta modell-ID i begÀran:

Inte tillgÀnglig via API:et

Robotikmodeller körs pÄ robotmaskinvara, inte via railwail-API:et.

05

Specifikationer

Modell-ID
rt-2-x
Utvecklare
Google DeepMind
Inmatning
Text, Bild
Utmatning
Roboteraktioner
Livscykel
Inte tillgÀnglig
Modellstorlek
Up to 55B (RT-2-X variants: 5B and 55B)
Licens
Research-only - Google DeepMind has not publicly released the RT-2-X weights, code or API. Some Open-X-Embodiment data and smaller RT-X reproductions are available, but the proprietary RT-2-X checkpoints are not.
KataloginlÀgg uppdaterat
23 september 2026

Taggar

  • google
  • vla
  • robotics
  • research-only
  • weights-closed
06

AnvÀndningsfall

Vad det anvÀnds till

  • Reference baseline in VLA / cross-embodiment papers
  • Comparative studies vs OpenVLA / Octo / π-0
  • Academic citation as the canonical large VLA
  • Demonstration of emergent reasoning in robotics
  • Motivating examples for action-token tokenisation research
  • Closed model - no direct deployment use
07

Vanliga frÄgor

Vad Àr Google RT-2-X?

Google RT-2-X Àr en modell av Google DeepMind i kategorin Robotik / VLA. Den finns i Railwail-katalogen men kan inte köras för nÀrvarande.

Vad kostar Google RT-2-X pÄ Railwail?

Google RT-2-X kan inte köras pÄ Railwail för nÀrvarande, sÄ det finns inget aktuellt pris. TillgÀngliga alternativ med priser visas lÀngre ned pÄ denna sida.

Hur snabb Àr Google RT-2-X?

Det finns Ànnu inte tillrÀckligt mÄnga uppmÀtta körningar av Google RT-2-X pÄ Railwail för att ange en körningstid. Det beror pÄ inmatningen, instÀllningarna och belastningen hos leverantören.

NÀr ska jag anvÀnda Google RT-2-X?

Google RT-2-X tillhör kategorin Robotik / VLA. Kategorisidan listar de andra modellerna av denna typ med deras priser.

Alla modeller: Robotik / VLA

Kan Google RT-2-X bearbeta bilder?

Ja. Google RT-2-X accepterar bilder som inmatning utöver text.

Kan jag anvÀnda Google RT-2-X just nu?

Inte tillgÀnglig för nÀrvarande. Sidan förblir online; tillgÀngliga alternativ frÄn samma kategori visas lÀngre ned.

Alla modeller via ett API

En API-nyckel för alla modeller pÄ Railwail. AnvÀndningen debiteras frÄn förbetald kredit, 1 kredit = 0,01 US$.