Octo Base

Robotik / VLANicht verfügbar
von OtherModell-ID: octo-base

Berkeley/Stanford 93M transformer diffusion policy. Pretrained on 800k Open-X-Embodiment episodes.

Status
Nicht verfügbar
Eingabe → Ausgabe
Text + Bild → Roboteraktionen
Entwickler
Other
Aktualisiert
23. September 2026

Octo Base ist derzeit nicht verfügbar

Die Angaben auf dieser Seite kannst du weiter nachlesen. Mit einer der verfügbaren Alternativen unten kannst du sofort ein vergleichbares Modell nutzen.

01

Playground

Octo Base

Forschungsmodell

Derzeit nicht verfügbar

Octo Base ist ein Robotik-Modell (Vision-Language-Action) und lässt sich nicht über die railwail-API ausführen.

02

Über Octo Base

Kurz gesagtStand: 23. September 2026

Octo Base ist ein Modell von Other aus der Kategorie Robotik / VLA. Über Railwail ist Octo Base derzeit nicht verfügbar.

Hintergrund

Über UC Berkeley / Stanford (Octo Model Team)

Gegründet 2023 · Berkeley & Stanford, California, USA

The Octo project is a collaboration of academic labs led by Sergey Levine (UC Berkeley BAIR) and Chelsea Finn (Stanford IRIS), with contributions from CMU, Google DeepMind, and Toyota Research Institute. Octo was first released in May 2024 alongside the Open-X-Embodiment dataset effort, with the goal of producing a generalist, fully open-source robot policy that any researcher can fine-tune on a new robot in hours. Octo introduced the recipe of a transformer policy with a diffusion action head trained on 800k cross-embodiment demonstrations, and it has become a de-facto baseline in academic VLA / generalist-policy research. The team released both Octo-Small (27M) and Octo-Base (93M) under Apache-2.0, alongside code, checkpoints and a fine-tuning toolkit.

UC Berkeley / Stanford (Octo Model Team) besuchen

Architektur

Transformer policy with diffusion action head (Vision-Language-Action)

Octo-Base is a transformer-based generalist robot policy. Inputs are tokenised RGB views and a natural-language instruction (encoded with a T5-base text encoder), interleaved with learnable readout tokens. The transformer trunk consumes this sequence and emits action latents that are decoded by a diffusion head producing continuous action chunks (default 4-step lookahead, 7-DoF end-effector deltas). The model was pretrained on roughly 800k demonstrations from 25 datasets in the Open-X-Embodiment collection, covering 9 robots, both single-arm and bimanual setups. Octo is intentionally embodiment-agnostic: action and proprioception spaces are encoded via shared adapters so the same backbone can be fine-tuned to new robots with as little as a few hundred demos. The diffusion head gives smooth, multimodal trajectories that outperform discrete-token VLAs on dexterous tasks at this scale.

Parameter
93M

Fähigkeiten

  • Generalist VLA policy across many robot embodiments
  • Trained on ~800k demos from Open-X-Embodiment
  • Diffusion action head produces smooth continuous actions
  • Natural-language instruction conditioning (T5 encoder)
  • Multi-view image inputs (primary + wrist cameras)
  • Designed for fast fine-tuning on new robots and tasks
  • Apache-2.0 open weights, code and recipes
  • Strong academic baseline for VLA papers
  • Best for: research, fine-tuning to new embodiments, generalist-policy benchmarks.

Training & Lizenz

~800,000 robot trajectories drawn from 25 Open-X-Embodiment-compatible datasets across 9 robot embodiments (Franka, WidowX, Bridge, RT-1 Everyday Robots, Berkeley UR5, etc.). Trained on TPU v4 / v5 hardware.

Lizenz: Apache-2.0 - fully open weights, code, and dataset references. Research-friendly; commercial use permitted under the licence.

Sicherheitstests: Octo is a research artifact; no formal red-teaming or RSP. Risks (unsafe robot motions) are mitigated at deployment time by application-specific safety controllers, not by the model itself.

Bekannte Grenzen

  • Modest 93M scale - underperforms 7B+ VLAs on hard generalisation
  • Optimised for 7-DoF end-effector control - bimanual humanoid action spaces need adapters
  • Limited language reasoning relative to LLM-backed VLAs
  • Image resolution capped (256x256)
  • Trained mostly on Western lab data - geographic bias
  • Long-horizon planning requires external prompt decomposition
03

Preise

Derzeit nicht verfügbar. Für dieses Modell gibt es derzeit keinen Preis, deshalb lässt es sich nicht ausführen.

04

API

Rufe Octo Base mit deinem Railwail-API-Schlüssel auf. Diese Modell-ID gehört in die Anfrage:

Nicht über die API verfügbar

Robotik-Modelle laufen auf Roboter-Hardware, nicht über die railwail-API.

05

Spezifikationen

Modell-ID
octo-base
Entwickler
Other
Kategorie
Robotik / VLA
Eingabe
Text, Bild
Ausgabe
Roboteraktionen
Modellgrösse
93M
Lizenz
Apache-2.0 - fully open weights, code, and dataset references. Research-friendly; commercial use permitted under the licence.
Katalogeintrag aktualisiert
23. September 2026

Schlagwörter

  • berkeley
  • stanford
  • vla
  • robotics
  • research-only
  • open-weights
  • small
06

Einsatzgebiete

Wofür es genutzt wird

  • Generalist policy fine-tuning for new robots
  • Academic research and ablations on VLAs
  • Baseline in cross-embodiment robot learning papers
  • Educational / coursework robotics projects
  • Bootstrapping data-efficient manipulation policies
  • Reproducible benchmarks on Open-X-Embodiment
07

Häufige Fragen

Was ist Octo Base?

Octo Base ist ein Modell von Other aus der Kategorie Robotik / VLA. Es steht im Railwail-Katalog, lässt sich derzeit aber nicht ausführen.

Was kostet Octo Base bei Railwail?

Octo Base lässt sich über Railwail derzeit nicht ausführen, deshalb gibt es keinen aktuellen Preis. Verfügbare Alternativen mit Preisen stehen weiter unten auf dieser Seite.

Wie schnell ist Octo Base?

Für Octo Base gibt es bei Railwail noch zu wenige gemessene Läufe, um eine Laufzeit anzugeben. Sie hängt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Wann ist Octo Base die richtige Wahl?

Octo Base gehört zur Kategorie Robotik / VLA. Die Kategorieseite listet die anderen Modelle dieser Art mit ihren Preisen.

Alle Modelle: Robotik / VLA

Kann Octo Base Bilder verarbeiten?

Ja. Octo Base nimmt neben Text auch Bilder als Eingabe an.

Kann ich Octo Base gerade nutzen?

Derzeit nicht verfügbar. Die Seite bleibt online; verfügbare Alternativen aus derselben Kategorie stehen weiter unten.

Alle Modelle über eine API

Ein API-Schlüssel für alle Modelle auf Railwail. Abgerechnet wird über vorab gekaufte Credits, 1 Credit = $ 0.01.