Model Family
RT-2
Google's Robotic Transformer 2 — a VLM fine-tuned to emit robot actions as text tokens.
Definition
RT-2 demonstrated that web-scale vision-language pre-training transfers to robot control: a fine-tuned PaLI-X / PaLM-E outputs discretised robot actions as text tokens. It established the vision-language-action paradigm that pi-zero and OpenVLA built on.
Common use cases
- Generalist manipulation
- Semantic grounding
- Sim-to-real