OpenVLA-7B

Робототехніка / VLAНедоступно
від OtherID моделі: openvla-7b

Stanford/Berkeley open VLA trained on 970k Open-X-Embodiment episodes. Supports LoRA fine-tuning.

Статус
Недоступно
Вхід → вихід
Текст + Зображення → Дії робота
Розробник
Other
Оновлено
23 вересня 2026 р.

OpenVLA-7B наразі недоступна

Ви можете прочитати деталі на цій сторінці. Виберіть одну з доступних альтернатив нижче, щоб одразу запустити порівняну модель.

01

Playground

OpenVLA-7B

Дослідницька модель

Наразі недоступна

OpenVLA-7B — це модель робототехніки (vision-language-action) і не може бути запущена через railwail API.

02

Про OpenVLA-7B

КороткоСтаном на 23 вересня 2026 р.

OpenVLA-7B — це модель від Other у категорії Робототехніка / VLA. На Railwail OpenVLA-7B наразі недоступна.

Фон

Про Stanford / UC Berkeley / Toyota Research Institute

Засновано 2024 · Stanford & Berkeley, California, USA

OpenVLA is the result of an academic-industry consortium led by Moo Jin Kim and colleagues at Stanford, UC Berkeley, and Toyota Research Institute (with contributors from MIT, Google DeepMind and Physical Intelligence). Released in June 2024, it was the first fully open-weights 7-billion-parameter Vision-Language-Action model trained on the Open-X-Embodiment dataset. OpenVLA was designed as a direct, reproducible, and parameter-efficient alternative to Google's closed RT-2 / RT-2-X, with the explicit goal of letting any lab fine-tune a 7B-class VLA on a single A100 / H100. The model, code, training recipe and fine-tuning toolkits (including LoRA) are all released under MIT-style permissive licences. OpenVLA quickly became a standard baseline in academic VLA research and the starting point for many downstream policies (CogACT, π-0-FAST baselines, embodied agent demos).

Відвідати Stanford / UC Berkeley / Toyota Research Institute

Архітектура

Vision-Language-Action (autoregressive discrete-token VLA)

OpenVLA combines a Llama-2-7B language backbone with a dual visual encoder that concatenates DINOv2 and SigLIP features, fused into the LLM via a Prismatic VLM-style projector. The model treats actions as discretised tokens: each continuous robot action dimension is binned into 256 bins, and the resulting tokens are appended to the LLM's vocabulary. Training is a single autoregressive next-token objective predicting both language and action tokens given image observations and a natural-language instruction. OpenVLA was trained on ~970k demonstration episodes from the Open-X-Embodiment dataset spanning 22+ robot embodiments and used Llama-2-7B as a pretrained text+code backbone, which the authors found markedly improves language grounding compared to scratch-trained VLAs. Parameter-efficient fine-tuning with LoRA is officially supported, making OpenVLA the de-facto open VLA workhorse.

Параметри
7B

Можливості

  • Fully open-weights 7B Vision-Language-Action model
  • Llama-2-7B backbone with DINOv2 + SigLIP vision
  • Discrete action-token decoding (256 bins per DoF)
  • Trained on ~970k Open-X-Embodiment episodes
  • LoRA fine-tuning officially supported
  • Strong language-grounded manipulation across robots
  • Fits on a single A100 / H100 with quantisation
  • MIT-style permissive licence on weights and code
  • Best for: research, reproducible VLA baselines, fine-tuning on new robots.

Навчання та ліцензія

~970,000 robot demonstration episodes from the Open-X-Embodiment dataset (RT-X collection), spanning 22+ robot embodiments and a wide range of manipulation tasks. Llama-2-7B and the DINOv2 + SigLIP vision encoders provide web-scale pretraining priors.

Ліцензія: MIT-style permissive licence on code and weights; Llama-2 components subject to Meta's Llama-2 Community Licence. Considered research-friendly open-weights.

Тестування безпеки: OpenVLA is a research artifact - no formal red-teaming. Safety in deployment is left to downstream systems (force limits, workspace constraints, e-stops).

Відомі обмеження

  • Discrete action tokens can limit smoothness vs diffusion policies
  • Inference latency on 7B is non-trivial for high-frequency control
  • Coverage skewed to Open-X tasks - novel embodiments need fine-tuning
  • Single-image / few-camera setup by default
  • English-only language conditioning
  • Llama-2 licence restrictions still apply to derived weights
03

Ціни

Наразі недоступно. Наразі для цієї моделі немає ціни, тому її не можна запустити.

04

API

Викличте OpenVLA-7B за допомогою вашого API-ключа Railwail. Використовуйте цей ID моделі в запиті:

Недоступно через API

Моделі робототехніки працюють на апаратному забезпеченні робота, а не через railwail API.

05

Характеристики

ID моделі
openvla-7b
Розробник
Other
Вхідні дані
Текст, Зображення
Вихідні дані
Дії робота
Розмір моделі
7B
Ліцензія
MIT-style permissive licence on code and weights; Llama-2 components subject to Meta's Llama-2 Community Licence. Considered research-friendly open-weights.
Запис у каталозі оновлено
23 вересня 2026 р.

Теги

  • stanford
  • berkeley
  • vla
  • robotics
  • research-only
  • open-weights
06

Варіанти використання

Для чого це використовується

  • Open VLA baseline for academic research
  • Fine-tuning to new robots via LoRA
  • Comparative studies vs RT-2 / π-0 / Octo
  • Language-conditioned manipulation research
  • Multi-task generalist policy training
  • Teaching VLA architecture and tokenisation
07

Часті питання

Що таке OpenVLA-7B?

OpenVLA-7B — це модель від Other у категорії Робототехніка / VLA. Вона внесена до каталогу Railwail, але наразі не може бути запущена.

Скільки коштує OpenVLA-7B на Railwail?

OpenVLA-7B наразі не можна запустити на Railwail, тому поточної ціни немає. Доступні альтернативи з цінами наведені далі на цій сторінці.

Наскільки швидкий OpenVLA-7B?

На Railwail поки що недостатньо виміряних запусків OpenVLA-7B, щоб вказати час виконання. Це залежить від введення, налаштувань та навантаження у постачальника.

Коли варто використовувати OpenVLA-7B?

OpenVLA-7B належить до категорії Робототехніка / VLA. На сторінці категорії наведені інші моделі цього типу з їхніми цінами.

Усі моделі: Робототехніка / VLA

Чи може OpenVLA-7B обробляти зображення?

Так. OpenVLA-7B приймає зображення як введення, крім тексту.

Чи можу я використовувати OpenVLA-7B прямо зараз?

Наразі недоступно. Сторінка залишається в мережі; доступні альтернативи з тієї ж категорії наведені далі.

Усі моделі через один API

Один API-ключ для всіх моделей на Railwail. Використання оплачується з попередньо поповнених кредитів, 1 кредит = 0,01 USD.