Concept
aka VLA

Vision-Language-Action

Foundation model class that maps images and language directly to robot actions.

Definition

VLA models extend vision-language pre-training with action outputs — joint angles, end-effector poses, gripper commands — usually discretised into the LLM's vocabulary. They are the leading approach to generalist robot policies (RT-2, OpenVLA, pi-zero, Octo).

Common use cases

  • Manipulation
  • Mobile robots
  • Bimanual control

Related terms

    Vision-Language-Action — AI Glossary | Railwail