Google RT-2-X

Robótica / VLAIndisponível
por Google DeepMindID do modelo: rt-2-x

Google's VLA from RT-X collaboration. Trained on Open-X-Embodiment (22 robots, 527 skills), positive transfer.

Status
Indisponível
Entrada → Saída
Texto + Imagem → Ações de robô
Desenvolvedor
Google DeepMind
Atualizado
23 de setembro de 2026

Google RT-2-X não está disponível no momento

Você ainda pode ler os detalhes nesta página. Escolha uma das alternativas disponíveis abaixo para executar um modelo comparável imediatamente.

01

Playground

Google RT-2-X

Modelo de pesquisa

Indisponível no momento

Google RT-2-X é um modelo de robótica (vision-language-action) e não pode ser executado através da API railwail.

02

Sobre Google RT-2-X

ResumoA partir de 23 de setembro de 2026

Google RT-2-X é um modelo de Google DeepMind na categoria Robótica / VLA. Google RT-2-X não está disponível no Railwail no momento.

Fundo

Sobre Google DeepMind

Fundado em 2010 · London, UK / Mountain View, USA

RT-2-X is Google DeepMind's flagship Robotic Transformer 2 (RT-2) model retrained on the Open-X-Embodiment dataset - the first large-scale, multi-institution effort to assemble a unified robot-learning dataset spanning many labs and robots. Open-X-Embodiment was organised in 2023 by Google DeepMind together with 21+ academic and industry institutions (Stanford, UC Berkeley, CMU, MIT, Toyota Research Institute, etc.), producing the RT-X dataset of ~1 million trajectories across 22 robot embodiments. RT-2-X extends RT-2's Vision-Language-Action recipe - using a PaLM-E / PaLI-X style VLM as backbone and emitting actions as text tokens - to this cross-embodiment corpus, demonstrating positive transfer across robots and a new state of the art on generalist manipulation at the time of release. RT-2-X is research-only and not publicly callable; it remains a key academic reference and the conceptual parent of subsequent open VLAs.

Visite Google DeepMind

Arquitetura

Vision-Language-Action transformer (PaLI / PaLM-E backbone, discrete action tokens)

RT-2-X follows the RT-2 design: a large Vision-Language Model (PaLI-X or PaLM-E) is co-fine-tuned on web-scale vision-language data and on robot demonstration data, where robot actions are tokenised as strings of natural-language-like tokens (each action dimension binned and rendered as a token). The same next-token prediction objective therefore trains the model on both internet-scale image-text data and on robot trajectories, allowing the resulting policy to inherit web knowledge (object semantics, OCR, common sense) and route it to motor commands. RT-2-X is the version of this recipe trained on the Open-X-Embodiment / RT-X dataset - ~1 million trajectories across 22 robot embodiments - rather than only on Google's internal kitchen-robot dataset. Public results report 5B and 55B variants, with the 55B model showing the strongest generalisation, especially when prompted with unseen language commands or unseen object combinations.

Parâmetros
Up to 55B (RT-2-X variants: 5B and 55B)

Capacidades

  • Generalist VLA trained on Open-X-Embodiment (22 robots)
  • Inherits web-scale knowledge from PaLI / PaLM-E backbones
  • Discrete action-token decoding (text-like vocabulary)
  • Positive transfer across robot embodiments
  • Strong emergent semantic reasoning (e.g. 'pick up the extinct animal')
  • 5B and 55B parameter variants
  • Reference architecture for the modern VLA paradigm
  • Co-training on internet data + robot demos
  • Best for: research, citation, conceptual baseline for VLAs.

Treinamento & licença

Co-trained on internet-scale vision-language data (PaLI / PaLM-E corpora) plus ~1 million robot trajectories from the Open-X-Embodiment (RT-X) dataset across 22 robot embodiments. Action targets are tokenised continuous controls.

Licença: Research-only - Google DeepMind has not publicly released the RT-2-X weights, code or API. Some Open-X-Embodiment data and smaller RT-X reproductions are available, but the proprietary RT-2-X checkpoints are not.

Testes de segurança: Covered under Google's standard Responsible AI process and frontier-safety evaluations. Robot demos were filmed in supervised lab settings with force-limited hardware; no public RSP-style document specific to RT-2-X.

Limitações conhecidas

  • Closed weights - no public API or download
  • Inference latency too high for very fast control loops
  • Discrete tokens limit smoothness vs diffusion / flow-matching policies
  • Cross-embodiment transfer still constrained by action-space differences
  • Long-horizon tasks need external prompt decomposition
  • Dataset skew toward kitchen / tabletop tasks
03

Preços

Atualmente indisponível. Não há preço para este modelo no momento, portanto não pode ser executado.

04

API

Chame Google RT-2-X com sua chave de API Railwail. Use este ID de modelo na solicitação:

Não disponível via API

Modelos de robótica são executados em hardware de robô, não através da API railwail.

05

Especificações

ID do modelo
rt-2-x
Desenvolvedor
Google DeepMind
Entrada
Texto, Imagem
Saída
Ações de robô
Ciclo de vida
Indisponível
Tamanho do modelo
Up to 55B (RT-2-X variants: 5B and 55B)
Licença
Research-only - Google DeepMind has not publicly released the RT-2-X weights, code or API. Some Open-X-Embodiment data and smaller RT-X reproductions are available, but the proprietary RT-2-X checkpoints are not.
Entrada do catálogo atualizada
23 de setembro de 2026

Etiquetas

  • google
  • vla
  • robotics
  • research-only
  • weights-closed
06

Casos de uso

Para que é utilizado

  • Reference baseline in VLA / cross-embodiment papers
  • Comparative studies vs OpenVLA / Octo / π-0
  • Academic citation as the canonical large VLA
  • Demonstration of emergent reasoning in robotics
  • Motivating examples for action-token tokenisation research
  • Closed model - no direct deployment use
07

Perguntas frequentes

O que é Google RT-2-X?

Google RT-2-X é um modelo de Google DeepMind na categoria Robótica / VLA. Está listado no Railwail, mas não pode ser executado no momento.

Quanto custa Google RT-2-X no Railwail?

Google RT-2-X não pode ser executado no Railwail no momento, portanto não há preço atual. Alternativas disponíveis com preços estão listadas mais abaixo nesta página.

Qual é a velocidade de Google RT-2-X?

Ainda não há execuções medidas suficientes de Google RT-2-X no Railwail para indicar um tempo de execução. Depende da entrada, das configurações e da carga no provedor.

Quando devo usar Google RT-2-X?

Google RT-2-X pertence à categoria Robótica / VLA. A página da categoria lista os outros modelos deste tipo com seus preços.

Todos os modelos: Robótica / VLA

Google RT-2-X pode processar imagens?

Sim. Google RT-2-X aceita imagens como entrada além de texto.

Posso usar Google RT-2-X agora?

Atualmente indisponível. A página permanece online; alternativas disponíveis da mesma categoria estão listadas mais abaixo.

Todos os modelos através de uma API

Uma chave API para todos os modelos no Railwail. O uso é cobrado a partir de créditos pré-pagos, 1 crédito = US$ 0,01.