OpenVLA-7B

Robotika / VLANedostupné
od OtherID modelu: openvla-7b

Stanford/Berkeley open VLA trained on 970k Open-X-Embodiment episodes. Supports LoRA fine-tuning.

Stav
Nedostupné
Vstup → výstup
Text + Obrázok → Akcie robota
Vývojár
Other
Aktualizované
23. septembra 2026

OpenVLA-7B nie je momentálne dostupný

Podrobnosti na tejto stránke si môžete prečítať. Vyberte si jednu z dostupných alternatív nižšie a spustite porovnateľný model hneď.

01

Playground

OpenVLA-7B

Výskumný model

Momentálne nedostupné

OpenVLA-7B je robotický model (vision-language-action) a nie je možné ho spustiť cez railwail API.

02

O OpenVLA-7B

StručneK 23. septembra 2026

OpenVLA-7B je model od Other v kategórii Robotika / VLA. OpenVLA-7B nie je v súčasnosti dostupný na Railwail.

Pozadie

O Stanford / UC Berkeley / Toyota Research Institute

Založené 2024 · Stanford & Berkeley, California, USA

OpenVLA is the result of an academic-industry consortium led by Moo Jin Kim and colleagues at Stanford, UC Berkeley, and Toyota Research Institute (with contributors from MIT, Google DeepMind and Physical Intelligence). Released in June 2024, it was the first fully open-weights 7-billion-parameter Vision-Language-Action model trained on the Open-X-Embodiment dataset. OpenVLA was designed as a direct, reproducible, and parameter-efficient alternative to Google's closed RT-2 / RT-2-X, with the explicit goal of letting any lab fine-tune a 7B-class VLA on a single A100 / H100. The model, code, training recipe and fine-tuning toolkits (including LoRA) are all released under MIT-style permissive licences. OpenVLA quickly became a standard baseline in academic VLA research and the starting point for many downstream policies (CogACT, π-0-FAST baselines, embodied agent demos).

Navštíviť Stanford / UC Berkeley / Toyota Research Institute

Architektúra

Vision-Language-Action (autoregressive discrete-token VLA)

OpenVLA combines a Llama-2-7B language backbone with a dual visual encoder that concatenates DINOv2 and SigLIP features, fused into the LLM via a Prismatic VLM-style projector. The model treats actions as discretised tokens: each continuous robot action dimension is binned into 256 bins, and the resulting tokens are appended to the LLM's vocabulary. Training is a single autoregressive next-token objective predicting both language and action tokens given image observations and a natural-language instruction. OpenVLA was trained on ~970k demonstration episodes from the Open-X-Embodiment dataset spanning 22+ robot embodiments and used Llama-2-7B as a pretrained text+code backbone, which the authors found markedly improves language grounding compared to scratch-trained VLAs. Parameter-efficient fine-tuning with LoRA is officially supported, making OpenVLA the de-facto open VLA workhorse.

Parametre
7B

Schopnosti

  • Fully open-weights 7B Vision-Language-Action model
  • Llama-2-7B backbone with DINOv2 + SigLIP vision
  • Discrete action-token decoding (256 bins per DoF)
  • Trained on ~970k Open-X-Embodiment episodes
  • LoRA fine-tuning officially supported
  • Strong language-grounded manipulation across robots
  • Fits on a single A100 / H100 with quantisation
  • MIT-style permissive licence on weights and code
  • Best for: research, reproducible VLA baselines, fine-tuning on new robots.

Tréning a licencia

~970,000 robot demonstration episodes from the Open-X-Embodiment dataset (RT-X collection), spanning 22+ robot embodiments and a wide range of manipulation tasks. Llama-2-7B and the DINOv2 + SigLIP vision encoders provide web-scale pretraining priors.

Licencia: MIT-style permissive licence on code and weights; Llama-2 components subject to Meta's Llama-2 Community Licence. Considered research-friendly open-weights.

Bezpečnostné testy: OpenVLA is a research artifact - no formal red-teaming. Safety in deployment is left to downstream systems (force limits, workspace constraints, e-stops).

Známe obmedzenia

  • Discrete action tokens can limit smoothness vs diffusion policies
  • Inference latency on 7B is non-trivial for high-frequency control
  • Coverage skewed to Open-X tasks - novel embodiments need fine-tuning
  • Single-image / few-camera setup by default
  • English-only language conditioning
  • Llama-2 licence restrictions still apply to derived weights
03

Ceny

Momentálne nedostupné. V súčasnosti nie je cena za tento model, preto ho nie je možné spustiť.

04

API

Zavolajte OpenVLA-7B s vaším API kľúčom Railwail. V požiadavke použite toto ID modelu:

Nie je dostupné cez API

Modely robotiky sa spúšťajú na hardvéri robota, nie cez railwail API.

05

Špecifikácie

ID modelu
openvla-7b
Vývojár
Other
Kategória
Robotika / VLA
Vstup
Text, Obrázok
Výstup
Akcie robota
Veľkosť modelu
7B
Licencia
MIT-style permissive licence on code and weights; Llama-2 components subject to Meta's Llama-2 Community Licence. Considered research-friendly open-weights.
Katalógová položka aktualizovaná
23. septembra 2026

Značky

  • stanford
  • berkeley
  • vla
  • robotics
  • research-only
  • open-weights
06

Prípady použitia

Na čo sa používa

  • Open VLA baseline for academic research
  • Fine-tuning to new robots via LoRA
  • Comparative studies vs RT-2 / π-0 / Octo
  • Language-conditioned manipulation research
  • Multi-task generalist policy training
  • Teaching VLA architecture and tokenisation
07

Často kladené otázky

Čo je OpenVLA-7B?

OpenVLA-7B je model od Other v kategórii Robotika / VLA. Je uvedený na Railwail, ale momentálne ho nie je možné spustiť.

Koľko stojí OpenVLA-7B na Railwail?

OpenVLA-7B nie je momentálne možné spustiť na Railwail, preto nie je aktuálna cena. Dostupné alternatívy s cenami sú uvedené nižšie na tejto stránke.

Ako rýchly je OpenVLA-7B?

Pre OpenVLA-7B je na Railwail zatiaľ príliš málo meraných spustení na určenie doby spustenia. Závisí to od vstupu, nastavení a zaťaženia u poskytovateľa.

Kedy by som mal používať OpenVLA-7B?

OpenVLA-7B patrí do kategórie Robotika / VLA. Stránka kategórie uvádza ďalšie modely tohto typu s ich cenami.

Všetky modely: Robotika / VLA

Môže OpenVLA-7B spracovávať obrázky?

Áno. OpenVLA-7B akceptuje obrázky ako vstup okrem textu.

Môžem OpenVLA-7B používať práve teraz?

Momentálne nedostupné. Stránka zostáva online; dostupné alternatívy z tej istej kategórie sú uvedené nižšie.

Všetky modely cez jedno API

Jeden API kľúč pre všetky modely na Railwail. Použitie sa účtuje z predplateného kreditu, 1 kredit = 0,01 USD.