Octo Base

Robotica / VLANon disponibile
di OtherID modello: octo-base

Berkeley/Stanford 93M transformer diffusion policy. Pretrained on 800k Open-X-Embodiment episodes.

Stato
Non disponibile
Input → output
Testo + Immagine → Azioni robot
Sviluppatore
Other
Aggiornato
23 settembre 2026

Octo Base non è attualmente disponibile

Puoi comunque leggere i dettagli su questa pagina. Scegli una delle alternative disponibili di seguito per eseguire subito un modello comparabile.

01

Playground

Octo Base

Modello di ricerca

Attualmente non disponibile

Octo Base è un modello di robotica (vision-language-action) e non può essere eseguito tramite l'API railwail.

02

Informazioni su Octo Base

RiassuntoA partire da 23 settembre 2026

Octo Base è un modello di Other nella categoria Robotica / VLA. Octo Base non è attualmente disponibile su Railwail.

Sfondo

Informazioni su UC Berkeley / Stanford (Octo Model Team)

Fondato 2023 · Berkeley & Stanford, California, USA

The Octo project is a collaboration of academic labs led by Sergey Levine (UC Berkeley BAIR) and Chelsea Finn (Stanford IRIS), with contributions from CMU, Google DeepMind, and Toyota Research Institute. Octo was first released in May 2024 alongside the Open-X-Embodiment dataset effort, with the goal of producing a generalist, fully open-source robot policy that any researcher can fine-tune on a new robot in hours. Octo introduced the recipe of a transformer policy with a diffusion action head trained on 800k cross-embodiment demonstrations, and it has become a de-facto baseline in academic VLA / generalist-policy research. The team released both Octo-Small (27M) and Octo-Base (93M) under Apache-2.0, alongside code, checkpoints and a fine-tuning toolkit.

Visita UC Berkeley / Stanford (Octo Model Team)

Architettura

Transformer policy with diffusion action head (Vision-Language-Action)

Octo-Base is a transformer-based generalist robot policy. Inputs are tokenised RGB views and a natural-language instruction (encoded with a T5-base text encoder), interleaved with learnable readout tokens. The transformer trunk consumes this sequence and emits action latents that are decoded by a diffusion head producing continuous action chunks (default 4-step lookahead, 7-DoF end-effector deltas). The model was pretrained on roughly 800k demonstrations from 25 datasets in the Open-X-Embodiment collection, covering 9 robots, both single-arm and bimanual setups. Octo is intentionally embodiment-agnostic: action and proprioception spaces are encoded via shared adapters so the same backbone can be fine-tuned to new robots with as little as a few hundred demos. The diffusion head gives smooth, multimodal trajectories that outperform discrete-token VLAs on dexterous tasks at this scale.

Parametri
93M

Capacità

  • Generalist VLA policy across many robot embodiments
  • Trained on ~800k demos from Open-X-Embodiment
  • Diffusion action head produces smooth continuous actions
  • Natural-language instruction conditioning (T5 encoder)
  • Multi-view image inputs (primary + wrist cameras)
  • Designed for fast fine-tuning on new robots and tasks
  • Apache-2.0 open weights, code and recipes
  • Strong academic baseline for VLA papers
  • Best for: research, fine-tuning to new embodiments, generalist-policy benchmarks.

Addestramento e licenza

~800,000 robot trajectories drawn from 25 Open-X-Embodiment-compatible datasets across 9 robot embodiments (Franka, WidowX, Bridge, RT-1 Everyday Robots, Berkeley UR5, etc.). Trained on TPU v4 / v5 hardware.

Licenza: Apache-2.0 - fully open weights, code, and dataset references. Research-friendly; commercial use permitted under the licence.

Test di sicurezza: Octo is a research artifact; no formal red-teaming or RSP. Risks (unsafe robot motions) are mitigated at deployment time by application-specific safety controllers, not by the model itself.

Limitazioni note

  • Modest 93M scale - underperforms 7B+ VLAs on hard generalisation
  • Optimised for 7-DoF end-effector control - bimanual humanoid action spaces need adapters
  • Limited language reasoning relative to LLM-backed VLAs
  • Image resolution capped (256x256)
  • Trained mostly on Western lab data - geographic bias
  • Long-horizon planning requires external prompt decomposition
03

Prezzi

Attualmente non disponibile. Al momento non c'è un prezzo per questo modello, quindi non può essere eseguito.

04

API

Chiama Octo Base con la tua chiave API Railwail. Usa questo ID modello nella richiesta:

Non disponibile tramite l'API

I modelli di robotica vengono eseguiti su hardware robotico, non tramite l'API railwail.

05

Specifiche

ID modello
octo-base
Sviluppatore
Other
Input
Testo, Immagine
Output
Azioni robot
Dimensione del modello
93M
Licenza
Apache-2.0 - fully open weights, code, and dataset references. Research-friendly; commercial use permitted under the licence.
Voce di catalogo aggiornata
23 settembre 2026

Etichette

  • berkeley
  • stanford
  • vla
  • robotics
  • research-only
  • open-weights
  • small
06

Casi d'uso

A cosa serve

  • Generalist policy fine-tuning for new robots
  • Academic research and ablations on VLAs
  • Baseline in cross-embodiment robot learning papers
  • Educational / coursework robotics projects
  • Bootstrapping data-efficient manipulation policies
  • Reproducible benchmarks on Open-X-Embodiment
07

Domande frequenti

Cos'è Octo Base?

Octo Base è un modello di Other nella categoria Robotica / VLA. È elencato su Railwail ma non può essere eseguito al momento.

Quanto costa Octo Base su Railwail?

Octo Base non può essere eseguito su Railwail al momento, quindi non c'è un prezzo attuale. Le alternative disponibili con i prezzi sono elencate più in basso in questa pagina.

Quanto è veloce Octo Base?

Non ci sono ancora abbastanza esecuzioni misurate di Octo Base su Railwail per indicare un tempo di esecuzione. Dipende dall'input, dalle impostazioni e dal carico presso il provider.

Quando devo usare Octo Base?

Octo Base appartiene alla categoria Robotica / VLA. La pagina della categoria elenca gli altri modelli di questo tipo con i loro prezzi.

Tutti i modelli: Robotica / VLA

Octo Base può elaborare immagini?

Sì. Octo Base accetta immagini come input oltre al testo.

Posso usare Octo Base adesso?

Attualmente non disponibile. La pagina rimane online; le alternative disponibili della stessa categoria sono elencate più in basso.

Tutti i modelli tramite un'API

Una chiave API per tutti i modelli su Railwail. L'utilizzo viene addebitato da crediti prepagati, 1 credito = 0,01 USD.