Google RT-2-X

Robotics / VLAUnavailable
by Google DeepMindModel ID: rt-2-x

Google's VLA from RT-X collaboration. Trained on Open-X-Embodiment (22 robots, 527 skills), positive transfer.

Status
Unavailable
Input โ†’ output
Text + Image โ†’ Robot actions
Developer
Google DeepMind
Updated
September 23, 2026

Google RT-2-X is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

01

Playground

Google RT-2-X

Research model

Currently unavailable

Google RT-2-X is a robotics model (vision-language-action) and cannot be run through the railwail API.

02

About Google RT-2-X

TL;DRAs of September 23, 2026

Google RT-2-X is a model by Google DeepMind in the Robotics / VLA category. Google RT-2-X is currently not available on Railwail.

Background

About Google DeepMind

Founded 2010 ยท London, UK / Mountain View, USA

RT-2-X is Google DeepMind's flagship Robotic Transformer 2 (RT-2) model retrained on the Open-X-Embodiment dataset - the first large-scale, multi-institution effort to assemble a unified robot-learning dataset spanning many labs and robots. Open-X-Embodiment was organised in 2023 by Google DeepMind together with 21+ academic and industry institutions (Stanford, UC Berkeley, CMU, MIT, Toyota Research Institute, etc.), producing the RT-X dataset of ~1 million trajectories across 22 robot embodiments. RT-2-X extends RT-2's Vision-Language-Action recipe - using a PaLM-E / PaLI-X style VLM as backbone and emitting actions as text tokens - to this cross-embodiment corpus, demonstrating positive transfer across robots and a new state of the art on generalist manipulation at the time of release. RT-2-X is research-only and not publicly callable; it remains a key academic reference and the conceptual parent of subsequent open VLAs.

Visit Google DeepMind

Architecture

Vision-Language-Action transformer (PaLI / PaLM-E backbone, discrete action tokens)

RT-2-X follows the RT-2 design: a large Vision-Language Model (PaLI-X or PaLM-E) is co-fine-tuned on web-scale vision-language data and on robot demonstration data, where robot actions are tokenised as strings of natural-language-like tokens (each action dimension binned and rendered as a token). The same next-token prediction objective therefore trains the model on both internet-scale image-text data and on robot trajectories, allowing the resulting policy to inherit web knowledge (object semantics, OCR, common sense) and route it to motor commands. RT-2-X is the version of this recipe trained on the Open-X-Embodiment / RT-X dataset - ~1 million trajectories across 22 robot embodiments - rather than only on Google's internal kitchen-robot dataset. Public results report 5B and 55B variants, with the 55B model showing the strongest generalisation, especially when prompted with unseen language commands or unseen object combinations.

Parameters
Up to 55B (RT-2-X variants: 5B and 55B)

Capabilities

  • Generalist VLA trained on Open-X-Embodiment (22 robots)
  • Inherits web-scale knowledge from PaLI / PaLM-E backbones
  • Discrete action-token decoding (text-like vocabulary)
  • Positive transfer across robot embodiments
  • Strong emergent semantic reasoning (e.g. 'pick up the extinct animal')
  • 5B and 55B parameter variants
  • Reference architecture for the modern VLA paradigm
  • Co-training on internet data + robot demos
  • Best for: research, citation, conceptual baseline for VLAs.

Training & license

Co-trained on internet-scale vision-language data (PaLI / PaLM-E corpora) plus ~1 million robot trajectories from the Open-X-Embodiment (RT-X) dataset across 22 robot embodiments. Action targets are tokenised continuous controls.

License: Research-only - Google DeepMind has not publicly released the RT-2-X weights, code or API. Some Open-X-Embodiment data and smaller RT-X reproductions are available, but the proprietary RT-2-X checkpoints are not.

Safety testing: Covered under Google's standard Responsible AI process and frontier-safety evaluations. Robot demos were filmed in supervised lab settings with force-limited hardware; no public RSP-style document specific to RT-2-X.

Known limitations

  • Closed weights - no public API or download
  • Inference latency too high for very fast control loops
  • Discrete tokens limit smoothness vs diffusion / flow-matching policies
  • Cross-embodiment transfer still constrained by action-space differences
  • Long-horizon tasks need external prompt decomposition
  • Dataset skew toward kitchen / tabletop tasks
03

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

04

API

Call Google RT-2-X with your Railwail API key. Use this model ID in the request:

Not available via the API

Robotics models run on robot hardware, not through the railwail API.

05

Specifications

Model ID
rt-2-x
Input
Text, Image
Output
Robot actions
Lifecycle
Unavailable
Model size
Up to 55B (RT-2-X variants: 5B and 55B)
License
Research-only - Google DeepMind has not publicly released the RT-2-X weights, code or API. Some Open-X-Embodiment data and smaller RT-X reproductions are available, but the proprietary RT-2-X checkpoints are not.
Catalog entry updated
September 23, 2026

Tags

  • google
  • vla
  • robotics
  • research-only
  • weights-closed
06

Use cases

What it is used for

  • Reference baseline in VLA / cross-embodiment papers
  • Comparative studies vs OpenVLA / Octo / ฯ€-0
  • Academic citation as the canonical large VLA
  • Demonstration of emergent reasoning in robotics
  • Motivating examples for action-token tokenisation research
  • Closed model - no direct deployment use
07

Frequently asked questions

What is Google RT-2-X?

Google RT-2-X is a model by Google DeepMind in the Robotics / VLA category. It is listed on Railwail but cannot be run at the moment.

How much does Google RT-2-X cost on Railwail?

Google RT-2-X cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

How fast is Google RT-2-X?

There are not enough measured runs of Google RT-2-X on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

When should I use Google RT-2-X?

Google RT-2-X belongs to the Robotics / VLA category. The category page lists the other models of this kind with their prices.

All models in Robotics / VLA

Can Google RT-2-X process images?

Yes. Google RT-2-X accepts images as input in addition to text.

Can I use Google RT-2-X right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.