OpenVLA-7B

Robotics / VLAUnavailable
by OtherModel ID: openvla-7b

Stanford/Berkeley open VLA trained on 970k Open-X-Embodiment episodes. Supports LoRA fine-tuning.

Status
Unavailable
Input โ†’ output
Text + Image โ†’ Robot actions
Developer
Other
Updated
September 23, 2026

OpenVLA-7B is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

01

Playground

OpenVLA-7B

Research model

Currently unavailable

OpenVLA-7B is a robotics model (vision-language-action) and cannot be run through the railwail API.

02

About OpenVLA-7B

TL;DRAs of September 23, 2026

OpenVLA-7B is a model by Other in the Robotics / VLA category. OpenVLA-7B is currently not available on Railwail.

Background

About Stanford / UC Berkeley / Toyota Research Institute

Founded 2024 ยท Stanford & Berkeley, California, USA

OpenVLA is the result of an academic-industry consortium led by Moo Jin Kim and colleagues at Stanford, UC Berkeley, and Toyota Research Institute (with contributors from MIT, Google DeepMind and Physical Intelligence). Released in June 2024, it was the first fully open-weights 7-billion-parameter Vision-Language-Action model trained on the Open-X-Embodiment dataset. OpenVLA was designed as a direct, reproducible, and parameter-efficient alternative to Google's closed RT-2 / RT-2-X, with the explicit goal of letting any lab fine-tune a 7B-class VLA on a single A100 / H100. The model, code, training recipe and fine-tuning toolkits (including LoRA) are all released under MIT-style permissive licences. OpenVLA quickly became a standard baseline in academic VLA research and the starting point for many downstream policies (CogACT, ฯ€-0-FAST baselines, embodied agent demos).

Visit Stanford / UC Berkeley / Toyota Research Institute

Architecture

Vision-Language-Action (autoregressive discrete-token VLA)

OpenVLA combines a Llama-2-7B language backbone with a dual visual encoder that concatenates DINOv2 and SigLIP features, fused into the LLM via a Prismatic VLM-style projector. The model treats actions as discretised tokens: each continuous robot action dimension is binned into 256 bins, and the resulting tokens are appended to the LLM's vocabulary. Training is a single autoregressive next-token objective predicting both language and action tokens given image observations and a natural-language instruction. OpenVLA was trained on ~970k demonstration episodes from the Open-X-Embodiment dataset spanning 22+ robot embodiments and used Llama-2-7B as a pretrained text+code backbone, which the authors found markedly improves language grounding compared to scratch-trained VLAs. Parameter-efficient fine-tuning with LoRA is officially supported, making OpenVLA the de-facto open VLA workhorse.

Parameters
7B

Capabilities

  • Fully open-weights 7B Vision-Language-Action model
  • Llama-2-7B backbone with DINOv2 + SigLIP vision
  • Discrete action-token decoding (256 bins per DoF)
  • Trained on ~970k Open-X-Embodiment episodes
  • LoRA fine-tuning officially supported
  • Strong language-grounded manipulation across robots
  • Fits on a single A100 / H100 with quantisation
  • MIT-style permissive licence on weights and code
  • Best for: research, reproducible VLA baselines, fine-tuning on new robots.

Training & license

~970,000 robot demonstration episodes from the Open-X-Embodiment dataset (RT-X collection), spanning 22+ robot embodiments and a wide range of manipulation tasks. Llama-2-7B and the DINOv2 + SigLIP vision encoders provide web-scale pretraining priors.

License: MIT-style permissive licence on code and weights; Llama-2 components subject to Meta's Llama-2 Community Licence. Considered research-friendly open-weights.

Safety testing: OpenVLA is a research artifact - no formal red-teaming. Safety in deployment is left to downstream systems (force limits, workspace constraints, e-stops).

Known limitations

  • Discrete action tokens can limit smoothness vs diffusion policies
  • Inference latency on 7B is non-trivial for high-frequency control
  • Coverage skewed to Open-X tasks - novel embodiments need fine-tuning
  • Single-image / few-camera setup by default
  • English-only language conditioning
  • Llama-2 licence restrictions still apply to derived weights
03

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

04

API

Call OpenVLA-7B with your Railwail API key. Use this model ID in the request:

Not available via the API

Robotics models run on robot hardware, not through the railwail API.

05

Specifications

Model ID
openvla-7b
Developer
Other
Input
Text, Image
Output
Robot actions
Model size
7B
License
MIT-style permissive licence on code and weights; Llama-2 components subject to Meta's Llama-2 Community Licence. Considered research-friendly open-weights.
Catalog entry updated
September 23, 2026

Tags

  • stanford
  • berkeley
  • vla
  • robotics
  • research-only
  • open-weights
06

Use cases

What it is used for

  • Open VLA baseline for academic research
  • Fine-tuning to new robots via LoRA
  • Comparative studies vs RT-2 / ฯ€-0 / Octo
  • Language-conditioned manipulation research
  • Multi-task generalist policy training
  • Teaching VLA architecture and tokenisation
07

Frequently asked questions

What is OpenVLA-7B?

OpenVLA-7B is a model by Other in the Robotics / VLA category. It is listed on Railwail but cannot be run at the moment.

How much does OpenVLA-7B cost on Railwail?

OpenVLA-7B cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

How fast is OpenVLA-7B?

There are not enough measured runs of OpenVLA-7B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

When should I use OpenVLA-7B?

OpenVLA-7B belongs to the Robotics / VLA category. The category page lists the other models of this kind with their prices.

All models in Robotics / VLA

Can OpenVLA-7B process images?

Yes. OpenVLA-7B accepts images as input in addition to text.

Can I use OpenVLA-7B right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.