LeRobot SmolVLA

Robotikk / VLAIkke tilgjengelig
av OtherModell-ID: smolvla

HuggingFace's 450M VLA pretrained on 487 community LeRobot datasets. Runs on consumer GPUs.

Status
Ikke tilgjengelig
Input → output
Tekst + Bilde → Roboteraksjoner
Utvikler
Other
Oppdatert
24. september 2026

LeRobot SmolVLA er for øyeblikket utilgjengelig

Du kan fortsatt lese detaljene på denne siden. Velg ett av de tilgjengelige alternativene nedenfor for å kjøre en sammenlignbar modell med en gang.

01

Playground

LeRobot SmolVLA

Forskningsmodell

Ikke tilgjengelig for øyeblikket

LeRobot SmolVLA er en robotikkmodell (vision-language-action) og kan ikke kjøres gjennom railwail API.

02

Om LeRobot SmolVLA

Kort sagtPer 24. september 2026

LeRobot SmolVLA er en modell fra Other i kategorien Robotikk / VLA. LeRobot SmolVLA er for øyeblikket ikke tilgjengelig på Railwail.

Bakgrunn

Om Hugging Face (LeRobot team)

Grunnlagt 2016 · New York, USA / Paris, France

SmolVLA is the flagship Vision-Language-Action model of Hugging Face's LeRobot project, an open-source robotics framework that brings the Transformers / Datasets philosophy to physical-AI research. SmolVLA was released in mid-2025 as a deliberately compact 450M-parameter VLA designed to be trainable and runnable on consumer hardware while still benefiting from community-scale pretraining. It is trained on 487 publicly contributed LeRobot community datasets - teleoperation episodes uploaded by hobbyists, university labs and small robotics companies - making it the first community-data-driven open VLA. The release includes pretraining and fine-tuning code, model checkpoints under Apache-2.0, and a tightly integrated stack with the LeRobot framework, hf-hub-hosted datasets, and the SO-100 / SO-ARM-100 low-cost robot arms.

Besøk Hugging Face (LeRobot team)

Arkitektur

Compact Vision-Language-Action transformer (action-chunk regression)

SmolVLA is a 450M-parameter transformer that combines a SmolVLM-style vision-language encoder with an action expert that regresses continuous action chunks. The vision-language tower is initialised from the open SmolVLM family (compact VLMs released by Hugging Face) and is responsible for fusing multi-view RGB observations with the natural-language instruction; a smaller action-prediction head consumes the resulting tokens together with proprioception and outputs a short chunk of continuous joint or end-effector actions. The model is pretrained on 487 LeRobot-format community datasets, covering single-arm, dual-arm and mobile-base setups, with a strong tilt toward the popular SO-100 and Koch low-cost teleoperation arms. Post-pretraining, users fine-tune on their own LeRobot recording for a specific robot and task. The whole stack is designed to run pretraining on a few H100s and fine-tuning on a single consumer GPU.

Parametere
450M

Evner

  • Compact 450M open VLA pretrained on community data
  • Trained on 487 LeRobot community datasets
  • SmolVLM-style vision-language tower + action expert
  • Continuous action-chunk regression
  • Runs fine-tuning on a single consumer GPU
  • Tight integration with LeRobot framework on Hugging Face
  • Apache-2.0 licence on weights and code
  • Strong baseline for SO-100 and Koch low-cost arms
  • Best for: hobbyists, educators, low-cost robot research.

Trening og lisens

487 publicly contributed LeRobot-format community datasets hosted on the Hugging Face Hub, dominated by teleoperation episodes from low-cost arms (SO-100, Koch) but also including dual-arm and mobile setups. Total scale on the order of millions of frames.

Lisens: Apache-2.0 - fully open weights, code, and datasets (where contributors used compatible licences). Designed for both research and commercial use.

Sikkerhetstesting: Research / hobbyist artifact - no formal red-teaming. Safety in deployment relies on low-torque hardware (SO-100, etc.) and user-supplied workspace constraints.

Kjente begrensninger

  • Modest scale - underperforms 7B VLAs on hard tasks
  • Dataset skew toward SO-100 / Koch low-cost arms
  • Limited language reasoning vs LLM-backed VLAs
  • Sensor coverage is mostly single RGB camera setups
  • Community data quality varies
  • Long-horizon behaviour limited without prompt decomposition
03

Priser

Ikke tilgjengelig for øyeblikket. Det er ingen pris for denne modellen for øyeblikket, så den kan ikke kjøres.

04

API

Ring LeRobot SmolVLA med din Railwail API-nøkkel. Bruk denne modell-IDen i forespørselen:

Ikke tilgjengelig via API-et

Robotikk-modeller kjører på robotmaskinvare, ikke gjennom railwail API.

05

Spesifikasjoner

Modell-ID
smolvla
Utvikler
Other
Inndata
Tekst, Bilde
Utdata
Roboteraksjoner
Modellstørrelse
450M
Lisens
Apache-2.0 - fully open weights, code, and datasets (where contributors used compatible licences). Designed for both research and commercial use.
Katalogoppføring oppdatert
24. september 2026

Merkelapper

  • huggingface
  • lerobot
  • vla
  • robotics
  • research-only
  • open-weights
  • small
  • consumer-gpu
06

Brukstilfeller

Hva det brukes til

  • Hobbyist robotics with low-cost arms (SO-100, Koch)
  • Educational coursework on VLAs
  • Community-data-driven robot learning research
  • Quick fine-tuning to new tasks on consumer GPUs
  • Reproducible baselines on LeRobot benchmarks
  • On-prem prototypes that need an Apache-2.0 VLA
07

Ofte stilte spørsmål

Hva er LeRobot SmolVLA?

LeRobot SmolVLA er en modell fra Other i kategorien Robotikk / VLA. Den er oppført på Railwail, men kan ikke kjøres for øyeblikket.

Hvor mye koster LeRobot SmolVLA på Railwail?

LeRobot SmolVLA kan ikke kjøres på Railwail for øyeblikket, så det finnes ingen gjeldende pris. Tilgjengelige alternativer med priser er oppført lenger ned på denne siden.

Hvor rask er LeRobot SmolVLA?

Det finnes ennå ikke nok målte kjøringer av LeRobot SmolVLA på Railwail til å angi en kjøretid. Det avhenger av inndataene, innstillingene og belastningen hos leverandøren.

Når bør jeg bruke LeRobot SmolVLA?

LeRobot SmolVLA tilhører kategorien Robotikk / VLA. Kategorisiden viser de andre modellene av denne typen med prisene deres.

Alle modeller: Robotikk / VLA

Kan LeRobot SmolVLA behandle bilder?

Ja. LeRobot SmolVLA godtar bilder som inndata i tillegg til tekst.

Kan jeg bruke LeRobot SmolVLA akkurat nå?

Ikke tilgjengelig for øyeblikket. Siden forblir online; tilgjengelige alternativer fra samme kategori er oppført lenger ned.

Alle modeller gjennom én API

Én API-nøkkel for alle modeller på Railwail. Bruk belastes fra forhåndsbetalt kreditt, 1 kreditt = 0,01 USD.