LeRobot SmolVLA

Robotik / VLANicht verfügbar
von OtherModell-ID: smolvla

HuggingFace's 450M VLA pretrained on 487 community LeRobot datasets. Runs on consumer GPUs.

Status
Nicht verfügbar
Eingabe → Ausgabe
Text + Bild → Roboteraktionen
Entwickler
Other
Aktualisiert
24. September 2026

LeRobot SmolVLA ist derzeit nicht verfügbar

Die Angaben auf dieser Seite kannst du weiter nachlesen. Mit einer der verfügbaren Alternativen unten kannst du sofort ein vergleichbares Modell nutzen.

01

Playground

LeRobot SmolVLA

Forschungsmodell

Derzeit nicht verfügbar

LeRobot SmolVLA ist ein Robotik-Modell (Vision-Language-Action) und lässt sich nicht über die railwail-API ausführen.

02

Über LeRobot SmolVLA

Kurz gesagtStand: 24. September 2026

LeRobot SmolVLA ist ein Modell von Other aus der Kategorie Robotik / VLA. Über Railwail ist LeRobot SmolVLA derzeit nicht verfügbar.

Hintergrund

Über Hugging Face (LeRobot team)

Gegründet 2016 · New York, USA / Paris, France

SmolVLA is the flagship Vision-Language-Action model of Hugging Face's LeRobot project, an open-source robotics framework that brings the Transformers / Datasets philosophy to physical-AI research. SmolVLA was released in mid-2025 as a deliberately compact 450M-parameter VLA designed to be trainable and runnable on consumer hardware while still benefiting from community-scale pretraining. It is trained on 487 publicly contributed LeRobot community datasets - teleoperation episodes uploaded by hobbyists, university labs and small robotics companies - making it the first community-data-driven open VLA. The release includes pretraining and fine-tuning code, model checkpoints under Apache-2.0, and a tightly integrated stack with the LeRobot framework, hf-hub-hosted datasets, and the SO-100 / SO-ARM-100 low-cost robot arms.

Hugging Face (LeRobot team) besuchen

Architektur

Compact Vision-Language-Action transformer (action-chunk regression)

SmolVLA is a 450M-parameter transformer that combines a SmolVLM-style vision-language encoder with an action expert that regresses continuous action chunks. The vision-language tower is initialised from the open SmolVLM family (compact VLMs released by Hugging Face) and is responsible for fusing multi-view RGB observations with the natural-language instruction; a smaller action-prediction head consumes the resulting tokens together with proprioception and outputs a short chunk of continuous joint or end-effector actions. The model is pretrained on 487 LeRobot-format community datasets, covering single-arm, dual-arm and mobile-base setups, with a strong tilt toward the popular SO-100 and Koch low-cost teleoperation arms. Post-pretraining, users fine-tune on their own LeRobot recording for a specific robot and task. The whole stack is designed to run pretraining on a few H100s and fine-tuning on a single consumer GPU.

Parameter
450M

Funktionen

  • Compact 450M open VLA pretrained on community data
  • Trained on 487 LeRobot community datasets
  • SmolVLM-style vision-language tower + action expert
  • Continuous action-chunk regression
  • Runs fine-tuning on a single consumer GPU
  • Tight integration with LeRobot framework on Hugging Face
  • Apache-2.0 licence on weights and code
  • Strong baseline for SO-100 and Koch low-cost arms
  • Best for: hobbyists, educators, low-cost robot research.

Training & Lizenz

487 publicly contributed LeRobot-format community datasets hosted on the Hugging Face Hub, dominated by teleoperation episodes from low-cost arms (SO-100, Koch) but also including dual-arm and mobile setups. Total scale on the order of millions of frames.

Lizenz: Apache-2.0 - fully open weights, code, and datasets (where contributors used compatible licences). Designed for both research and commercial use.

Sicherheitstests: Research / hobbyist artifact - no formal red-teaming. Safety in deployment relies on low-torque hardware (SO-100, etc.) and user-supplied workspace constraints.

Bekannte Einschränkungen

  • Modest scale - underperforms 7B VLAs on hard tasks
  • Dataset skew toward SO-100 / Koch low-cost arms
  • Limited language reasoning vs LLM-backed VLAs
  • Sensor coverage is mostly single RGB camera setups
  • Community data quality varies
  • Long-horizon behaviour limited without prompt decomposition
03

Preise

Derzeit nicht verfügbar. Für dieses Modell gibt es derzeit keinen Preis, deshalb lässt es sich nicht ausführen.

04

API

Rufe LeRobot SmolVLA mit deinem Railwail-API-Schlüssel auf. Diese Modell-ID gehört in die Anfrage:

Nicht über die API verfügbar

Robotik-Modelle laufen auf Roboter-Hardware, nicht über die railwail-API.

05

Spezifikationen

Modell-ID
smolvla
Entwickler
Other
Kategorie
Robotik / VLA
Eingabe
Text, Bild
Ausgabe
Roboteraktionen
Modellgröße
450M
Lizenz
Apache-2.0 - fully open weights, code, and datasets (where contributors used compatible licences). Designed for both research and commercial use.
Katalogeintrag aktualisiert
24. September 2026

Schlagwörter

  • huggingface
  • lerobot
  • vla
  • robotics
  • research-only
  • open-weights
  • small
  • consumer-gpu
06

Einsatzgebiete

Wofür es genutzt wird

  • Hobbyist robotics with low-cost arms (SO-100, Koch)
  • Educational coursework on VLAs
  • Community-data-driven robot learning research
  • Quick fine-tuning to new tasks on consumer GPUs
  • Reproducible baselines on LeRobot benchmarks
  • On-prem prototypes that need an Apache-2.0 VLA
07

Häufige Fragen

Was ist LeRobot SmolVLA?

LeRobot SmolVLA ist ein Modell von Other aus der Kategorie Robotik / VLA. Es steht im Railwail-Katalog, lässt sich derzeit aber nicht ausführen.

Was kostet LeRobot SmolVLA bei Railwail?

LeRobot SmolVLA lässt sich über Railwail derzeit nicht ausführen, deshalb gibt es keinen aktuellen Preis. Verfügbare Alternativen mit Preisen stehen weiter unten auf dieser Seite.

Wie schnell ist LeRobot SmolVLA?

Für LeRobot SmolVLA gibt es bei Railwail noch zu wenige gemessene Läufe, um eine Laufzeit anzugeben. Sie hängt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Wann sollte ich LeRobot SmolVLA nutzen?

LeRobot SmolVLA gehört zur Kategorie Robotik / VLA. Die Kategorieseite listet die anderen Modelle dieser Art mit ihren Preisen.

Alle Modelle: Robotik / VLA

Kann LeRobot SmolVLA Bilder verarbeiten?

Ja. LeRobot SmolVLA nimmt neben Text auch Bilder als Eingabe an.

Kann ich LeRobot SmolVLA gerade nutzen?

Derzeit nicht verfügbar. Die Seite bleibt online; verfügbare Alternativen aus derselben Kategorie stehen weiter unten.

Alle Modelle über eine API

Ein API-Schlüssel für alle Modelle auf Railwail. Abgerechnet wird über vorab gekaufte Credits, 1 Credit = $ 0,01.