Octo Base

Robotik / VLAIkke tilgængelig
af OtherModell-ID: octo-base

Berkeley/Stanford 93M transformer diffusion policy. Pretrained on 800k Open-X-Embodiment episodes.

Status
Ikke tilgængelig
Input → output
Tekst + Billede → Roboteraktioner
Udvikler
Other
Opdateret
23. september 2026

Octo Base er i øjeblikket utilgængelig

Du kan stadig læse detaljerne på denne side. Vælg en af de tilgængelige alternativer nedenfor for at køre en sammenlignelig model med det samme.

01

Playground

Octo Base

Forskningsmodel

Derzeit nicht verfügbar

Octo Base er en robotik-model (vision-language-action) og kan ikke køres via railwail API.

02

Om Octo Base

Kort sagtFra 23. september 2026

Octo Base er en model af Other i kategorien Robotik / VLA. Octo Base er i øjeblikket ikke tilgængelig på Railwail.

Baggrund

Om UC Berkeley / Stanford (Octo Model Team)

Grundlagt 2023 · Berkeley & Stanford, California, USA

The Octo project is a collaboration of academic labs led by Sergey Levine (UC Berkeley BAIR) and Chelsea Finn (Stanford IRIS), with contributions from CMU, Google DeepMind, and Toyota Research Institute. Octo was first released in May 2024 alongside the Open-X-Embodiment dataset effort, with the goal of producing a generalist, fully open-source robot policy that any researcher can fine-tune on a new robot in hours. Octo introduced the recipe of a transformer policy with a diffusion action head trained on 800k cross-embodiment demonstrations, and it has become a de-facto baseline in academic VLA / generalist-policy research. The team released both Octo-Small (27M) and Octo-Base (93M) under Apache-2.0, alongside code, checkpoints and a fine-tuning toolkit.

Besøg UC Berkeley / Stanford (Octo Model Team)

Arkitektur

Transformer policy with diffusion action head (Vision-Language-Action)

Octo-Base is a transformer-based generalist robot policy. Inputs are tokenised RGB views and a natural-language instruction (encoded with a T5-base text encoder), interleaved with learnable readout tokens. The transformer trunk consumes this sequence and emits action latents that are decoded by a diffusion head producing continuous action chunks (default 4-step lookahead, 7-DoF end-effector deltas). The model was pretrained on roughly 800k demonstrations from 25 datasets in the Open-X-Embodiment collection, covering 9 robots, both single-arm and bimanual setups. Octo is intentionally embodiment-agnostic: action and proprioception spaces are encoded via shared adapters so the same backbone can be fine-tuned to new robots with as little as a few hundred demos. The diffusion head gives smooth, multimodal trajectories that outperform discrete-token VLAs on dexterous tasks at this scale.

Parametre
93M

Funktioner

  • Generalist VLA policy across many robot embodiments
  • Trained on ~800k demos from Open-X-Embodiment
  • Diffusion action head produces smooth continuous actions
  • Natural-language instruction conditioning (T5 encoder)
  • Multi-view image inputs (primary + wrist cameras)
  • Designed for fast fine-tuning on new robots and tasks
  • Apache-2.0 open weights, code and recipes
  • Strong academic baseline for VLA papers
  • Best for: research, fine-tuning to new embodiments, generalist-policy benchmarks.

Træning og licens

~800,000 robot trajectories drawn from 25 Open-X-Embodiment-compatible datasets across 9 robot embodiments (Franka, WidowX, Bridge, RT-1 Everyday Robots, Berkeley UR5, etc.). Trained on TPU v4 / v5 hardware.

Licens: Apache-2.0 - fully open weights, code, and dataset references. Research-friendly; commercial use permitted under the licence.

Sikkerhedstests: Octo is a research artifact; no formal red-teaming or RSP. Risks (unsafe robot motions) are mitigated at deployment time by application-specific safety controllers, not by the model itself.

Kendte begrænsninger

  • Modest 93M scale - underperforms 7B+ VLAs on hard generalisation
  • Optimised for 7-DoF end-effector control - bimanual humanoid action spaces need adapters
  • Limited language reasoning relative to LLM-backed VLAs
  • Image resolution capped (256x256)
  • Trained mostly on Western lab data - geographic bias
  • Long-horizon planning requires external prompt decomposition
03

Priser

Ikke tilgængelig i øjeblikket. Der er i øjeblikket ingen pris for denne model, så den kan ikke køres.

04

API

Kald Octo Base med din Railwail API-nøgle. Brug dette model-ID i anmodningen:

Ikke tilgængelig via API'en

Robotik-modeller køres på robotik-hardware, ikke gennem railwail API'en.

05

Specifikationer

Model-ID
octo-base
Udvikler
Other
Input
Tekst, Billede
Output
Roboteraktioner
Modelstørrelse
93M
Licens
Apache-2.0 - fully open weights, code, and dataset references. Research-friendly; commercial use permitted under the licence.
Katalogelement opdateret
23. september 2026

Tags

  • berkeley
  • stanford
  • vla
  • robotics
  • research-only
  • open-weights
  • small
06

Anvendelsestilfælde

Hvad det bruges til

  • Generalist policy fine-tuning for new robots
  • Academic research and ablations on VLAs
  • Baseline in cross-embodiment robot learning papers
  • Educational / coursework robotics projects
  • Bootstrapping data-efficient manipulation policies
  • Reproducible benchmarks on Open-X-Embodiment
07

Ofte stillede spørgsmål

Hvad er Octo Base?

Octo Base er en model fra Other i kategorien Robotik / VLA. Den er opført på Railwail, men kan ikke køres i øjeblikket.

Hvad koster Octo Base på Railwail?

Octo Base kan ikke køres på Railwail i øjeblikket, så der er ingen aktuel pris. Tilgængelige alternativer med priser er angivet længere nede på denne side.

Hvor hurtig er Octo Base?

Der er endnu ikke nok målte kørsler af Octo Base på Railwail til at angive en udførelsestid. Det afhænger af inputtet, indstillingerne og belastningen hos provideren.

Hvornår skal jeg bruge Octo Base?

Octo Base tilhører kategorien Robotik / VLA. Kategorisiden viser de andre modeller af denne type med deres priser.

Alle modeller: Robotik / VLA

Kan Octo Base behandle billeder?

Ja. Octo Base accepterer billeder som input ud over tekst.

Kan jeg bruge Octo Base lige nu?

Ikke tilgængelig i øjeblikket. Siden forbliver online; tilgængelige alternativer fra samme kategori er angivet længere nede.

Alle modeller via én API

En API-nøgle til alle modeller på Railwail. Forbrug debiteres fra forudbetalte credits, 1 credit = 0,01 US$.