OpenVLA-7B

Robotik / VLANicht verfügbar
von OtherModell-ID: openvla-7b

Stanford/Berkeley open VLA trained on 970k Open-X-Embodiment episodes. Supports LoRA fine-tuning.

Status
Nicht verfügbar
Eingabe → Ausgabe
Text + Bild → Roboteraktionen
Entwickler
Other
Aktualisiert
23. September 2026

OpenVLA-7B ist derzeit nicht verfügbar

Die Angaben auf dieser Seite kannst du weiter nachlesen. Mit einer der verfügbaren Alternativen unten kannst du sofort ein vergleichbares Modell nutzen.

01

Playground

OpenVLA-7B

Forschungsmodell

Derzeit nicht verfügbar

OpenVLA-7B ist ein Robotik-Modell (Vision-Language-Action) und lässt sich nicht über die railwail-API ausführen.

02

Über OpenVLA-7B

Kurz gesagtStand: 23. September 2026

OpenVLA-7B ist ein Modell von Other aus der Kategorie Robotik / VLA. Über Railwail ist OpenVLA-7B derzeit nicht verfügbar.

Hintergrund

Über Stanford / UC Berkeley / Toyota Research Institute

Gegründet 2024 · Stanford & Berkeley, California, USA

OpenVLA is the result of an academic-industry consortium led by Moo Jin Kim and colleagues at Stanford, UC Berkeley, and Toyota Research Institute (with contributors from MIT, Google DeepMind and Physical Intelligence). Released in June 2024, it was the first fully open-weights 7-billion-parameter Vision-Language-Action model trained on the Open-X-Embodiment dataset. OpenVLA was designed as a direct, reproducible, and parameter-efficient alternative to Google's closed RT-2 / RT-2-X, with the explicit goal of letting any lab fine-tune a 7B-class VLA on a single A100 / H100. The model, code, training recipe and fine-tuning toolkits (including LoRA) are all released under MIT-style permissive licences. OpenVLA quickly became a standard baseline in academic VLA research and the starting point for many downstream policies (CogACT, π-0-FAST baselines, embodied agent demos).

Stanford / UC Berkeley / Toyota Research Institute besuchen

Architektur

Vision-Language-Action (autoregressive discrete-token VLA)

OpenVLA combines a Llama-2-7B language backbone with a dual visual encoder that concatenates DINOv2 and SigLIP features, fused into the LLM via a Prismatic VLM-style projector. The model treats actions as discretised tokens: each continuous robot action dimension is binned into 256 bins, and the resulting tokens are appended to the LLM's vocabulary. Training is a single autoregressive next-token objective predicting both language and action tokens given image observations and a natural-language instruction. OpenVLA was trained on ~970k demonstration episodes from the Open-X-Embodiment dataset spanning 22+ robot embodiments and used Llama-2-7B as a pretrained text+code backbone, which the authors found markedly improves language grounding compared to scratch-trained VLAs. Parameter-efficient fine-tuning with LoRA is officially supported, making OpenVLA the de-facto open VLA workhorse.

Parameter
7B

Funktionen

  • Fully open-weights 7B Vision-Language-Action model
  • Llama-2-7B backbone with DINOv2 + SigLIP vision
  • Discrete action-token decoding (256 bins per DoF)
  • Trained on ~970k Open-X-Embodiment episodes
  • LoRA fine-tuning officially supported
  • Strong language-grounded manipulation across robots
  • Fits on a single A100 / H100 with quantisation
  • MIT-style permissive licence on weights and code
  • Best for: research, reproducible VLA baselines, fine-tuning on new robots.

Training & Lizenz

~970,000 robot demonstration episodes from the Open-X-Embodiment dataset (RT-X collection), spanning 22+ robot embodiments and a wide range of manipulation tasks. Llama-2-7B and the DINOv2 + SigLIP vision encoders provide web-scale pretraining priors.

Lizenz: MIT-style permissive licence on code and weights; Llama-2 components subject to Meta's Llama-2 Community Licence. Considered research-friendly open-weights.

Sicherheitstests: OpenVLA is a research artifact - no formal red-teaming. Safety in deployment is left to downstream systems (force limits, workspace constraints, e-stops).

Bekannte Einschränkungen

  • Discrete action tokens can limit smoothness vs diffusion policies
  • Inference latency on 7B is non-trivial for high-frequency control
  • Coverage skewed to Open-X tasks - novel embodiments need fine-tuning
  • Single-image / few-camera setup by default
  • English-only language conditioning
  • Llama-2 licence restrictions still apply to derived weights
03

Preise

Derzeit nicht verfügbar. Für dieses Modell gibt es derzeit keinen Preis, deshalb lässt es sich nicht ausführen.

04

API

Rufe OpenVLA-7B mit deinem Railwail-API-Schlüssel auf. Diese Modell-ID gehört in die Anfrage:

Nicht über die API verfügbar

Robotik-Modelle laufen auf Roboter-Hardware, nicht über die railwail-API.

05

Spezifikationen

Modell-ID
openvla-7b
Entwickler
Other
Kategorie
Robotik / VLA
Eingabe
Text, Bild
Ausgabe
Roboteraktionen
Modellgröße
7B
Lizenz
MIT-style permissive licence on code and weights; Llama-2 components subject to Meta's Llama-2 Community Licence. Considered research-friendly open-weights.
Katalogeintrag aktualisiert
23. September 2026

Schlagwörter

  • stanford
  • berkeley
  • vla
  • robotics
  • research-only
  • open-weights
06

Einsatzgebiete

Wofür es genutzt wird

  • Open VLA baseline for academic research
  • Fine-tuning to new robots via LoRA
  • Comparative studies vs RT-2 / π-0 / Octo
  • Language-conditioned manipulation research
  • Multi-task generalist policy training
  • Teaching VLA architecture and tokenisation
07

Häufige Fragen

Was ist OpenVLA-7B?

OpenVLA-7B ist ein Modell von Other aus der Kategorie Robotik / VLA. Es steht im Railwail-Katalog, lässt sich derzeit aber nicht ausführen.

Was kostet OpenVLA-7B bei Railwail?

OpenVLA-7B lässt sich über Railwail derzeit nicht ausführen, deshalb gibt es keinen aktuellen Preis. Verfügbare Alternativen mit Preisen stehen weiter unten auf dieser Seite.

Wie schnell ist OpenVLA-7B?

Für OpenVLA-7B gibt es bei Railwail noch zu wenige gemessene Läufe, um eine Laufzeit anzugeben. Sie hängt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Wann sollte ich OpenVLA-7B nutzen?

OpenVLA-7B gehört zur Kategorie Robotik / VLA. Die Kategorieseite listet die anderen Modelle dieser Art mit ihren Preisen.

Alle Modelle: Robotik / VLA

Kann OpenVLA-7B Bilder verarbeiten?

Ja. OpenVLA-7B nimmt neben Text auch Bilder als Eingabe an.

Kann ich OpenVLA-7B gerade nutzen?

Derzeit nicht verfügbar. Die Seite bleibt online; verfügbare Alternativen aus derselben Kategorie stehen weiter unten.

Alle Modelle über eine API

Ein API-Schlüssel für alle Modelle auf Railwail. Abgerechnet wird über vorab gekaufte Credits, 1 Credit = $ 0,01.