Google RT-2-X

Robótica / VLANo disponible
de Google DeepMindID del modelo: rt-2-x

Google's VLA from RT-X collaboration. Trained on Open-X-Embodiment (22 robots, 527 skills), positive transfer.

Estado
No disponible
Entrada → Salida
Texto + Imagen → Acciones de robot
Desarrollador
Google DeepMind
Actualizado
23 de septiembre de 2026

Google RT-2-X no está disponible actualmente

Puedes seguir leyendo los detalles en esta página. Elige una de las alternativas disponibles abajo para ejecutar un modelo comparable de inmediato.

01

Playground

Google RT-2-X

Modelo de investigación

No disponible actualmente

Google RT-2-X es un modelo de robótica (visión-lenguaje-acción) y no se puede ejecutar a través de la API de railwail.

02

Acerca de Google RT-2-X

ResumenA fecha de 23 de septiembre de 2026

Google RT-2-X es un modelo de Google DeepMind en la categoría Robótica / VLA. Google RT-2-X no está disponible en Railwail en este momento.

Fondo

Acerca de Google DeepMind

Fundado en 2010 · London, UK / Mountain View, USA

RT-2-X is Google DeepMind's flagship Robotic Transformer 2 (RT-2) model retrained on the Open-X-Embodiment dataset - the first large-scale, multi-institution effort to assemble a unified robot-learning dataset spanning many labs and robots. Open-X-Embodiment was organised in 2023 by Google DeepMind together with 21+ academic and industry institutions (Stanford, UC Berkeley, CMU, MIT, Toyota Research Institute, etc.), producing the RT-X dataset of ~1 million trajectories across 22 robot embodiments. RT-2-X extends RT-2's Vision-Language-Action recipe - using a PaLM-E / PaLI-X style VLM as backbone and emitting actions as text tokens - to this cross-embodiment corpus, demonstrating positive transfer across robots and a new state of the art on generalist manipulation at the time of release. RT-2-X is research-only and not publicly callable; it remains a key academic reference and the conceptual parent of subsequent open VLAs.

Visitar Google DeepMind

Arquitectura

Vision-Language-Action transformer (PaLI / PaLM-E backbone, discrete action tokens)

RT-2-X follows the RT-2 design: a large Vision-Language Model (PaLI-X or PaLM-E) is co-fine-tuned on web-scale vision-language data and on robot demonstration data, where robot actions are tokenised as strings of natural-language-like tokens (each action dimension binned and rendered as a token). The same next-token prediction objective therefore trains the model on both internet-scale image-text data and on robot trajectories, allowing the resulting policy to inherit web knowledge (object semantics, OCR, common sense) and route it to motor commands. RT-2-X is the version of this recipe trained on the Open-X-Embodiment / RT-X dataset - ~1 million trajectories across 22 robot embodiments - rather than only on Google's internal kitchen-robot dataset. Public results report 5B and 55B variants, with the 55B model showing the strongest generalisation, especially when prompted with unseen language commands or unseen object combinations.

Parámetros
Up to 55B (RT-2-X variants: 5B and 55B)

Capacidades

  • Generalist VLA trained on Open-X-Embodiment (22 robots)
  • Inherits web-scale knowledge from PaLI / PaLM-E backbones
  • Discrete action-token decoding (text-like vocabulary)
  • Positive transfer across robot embodiments
  • Strong emergent semantic reasoning (e.g. 'pick up the extinct animal')
  • 5B and 55B parameter variants
  • Reference architecture for the modern VLA paradigm
  • Co-training on internet data + robot demos
  • Best for: research, citation, conceptual baseline for VLAs.

Entrenamiento y licencia

Co-trained on internet-scale vision-language data (PaLI / PaLM-E corpora) plus ~1 million robot trajectories from the Open-X-Embodiment (RT-X) dataset across 22 robot embodiments. Action targets are tokenised continuous controls.

Licencia: Research-only - Google DeepMind has not publicly released the RT-2-X weights, code or API. Some Open-X-Embodiment data and smaller RT-X reproductions are available, but the proprietary RT-2-X checkpoints are not.

Pruebas de seguridad: Covered under Google's standard Responsible AI process and frontier-safety evaluations. Robot demos were filmed in supervised lab settings with force-limited hardware; no public RSP-style document specific to RT-2-X.

Limitaciones conocidas

  • Closed weights - no public API or download
  • Inference latency too high for very fast control loops
  • Discrete tokens limit smoothness vs diffusion / flow-matching policies
  • Cross-embodiment transfer still constrained by action-space differences
  • Long-horizon tasks need external prompt decomposition
  • Dataset skew toward kitchen / tabletop tasks
03

Precios

Actualmente no disponible. No hay precio para este modelo en este momento, por lo que no se puede ejecutar.

04

API

Llama a Google RT-2-X con tu clave API de Railwail. Usa este ID de modelo en la solicitud:

No disponible a través de la API

Los modelos de robótica se ejecutan en hardware de robot, no a través de la API de railwail.

05

Especificaciones

ID del modelo
rt-2-x
Desarrollador
Google DeepMind
Categoría
Robótica / VLA
Entrada
Texto, Imagen
Salida
Acciones de robot
Ciclo de vida
No disponible
Tamaño del modelo
Up to 55B (RT-2-X variants: 5B and 55B)
Licencia
Research-only - Google DeepMind has not publicly released the RT-2-X weights, code or API. Some Open-X-Embodiment data and smaller RT-X reproductions are available, but the proprietary RT-2-X checkpoints are not.
Entrada del catálogo actualizada
23 de septiembre de 2026

Etiquetas

  • google
  • vla
  • robotics
  • research-only
  • weights-closed
06

Casos de uso

Para qué se utiliza

  • Reference baseline in VLA / cross-embodiment papers
  • Comparative studies vs OpenVLA / Octo / π-0
  • Academic citation as the canonical large VLA
  • Demonstration of emergent reasoning in robotics
  • Motivating examples for action-token tokenisation research
  • Closed model - no direct deployment use
07

Preguntas frecuentes

¿Qué es Google RT-2-X?

Google RT-2-X es un modelo de Google DeepMind en la categoría Robótica / VLA. Aparece en el catálogo de Railwail pero no se puede ejecutar en este momento.

¿Cuánto cuesta Google RT-2-X en Railwail?

Google RT-2-X no se puede ejecutar en Railwail en este momento, por lo que no hay precio actual. Las alternativas disponibles con precios se enumeran más abajo en esta página.

¿Qué velocidad tiene Google RT-2-X?

Aún no hay suficientes ejecuciones medidas de Google RT-2-X en Railwail para indicar un tiempo de ejecución. Depende de la entrada, la configuración y la carga en el proveedor.

¿Cuándo debo usar Google RT-2-X?

Google RT-2-X pertenece a la categoría Robótica / VLA. La página de categoría enumera los otros modelos de este tipo con sus precios.

Todos los modelos: Robótica / VLA

¿Puede Google RT-2-X procesar imágenes?

Sí. Google RT-2-X acepta imágenes como entrada además de texto.

¿Puedo usar Google RT-2-X ahora mismo?

Actualmente no disponible. La página se mantiene en línea; las alternativas disponibles de la misma categoría se enumeran más abajo.

Todos los modelos a través de una API

Una clave API para todos los modelos en Railwail. El uso se cobra desde créditos prepagados, 1 crédito = 0,01 US$.