Concept
aka VLA
Vision-Language-Action
Foundation model class that maps images and language directly to robot actions.
Definition
VLA models extend vision-language pre-training with action outputs — joint angles, end-effector poses, gripper commands — usually discretised into the LLM's vocabulary. They are the leading approach to generalist robot policies (RT-2, OpenVLA, pi-zero, Octo).
Common use cases
- Manipulation
- Mobile robots
- Bimanual control