Octo Base

Robotics / VLAUnavailable
by OtherModel ID: octo-base

Berkeley/Stanford 93M transformer diffusion policy. Pretrained on 800k Open-X-Embodiment episodes.

Status
Unavailable
Input โ†’ output
Text + Image โ†’ Robot actions
Developer
Other
Updated
September 23, 2026

Octo Base is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

01

Playground

Octo Base

Research model

Currently unavailable

Octo Base is a robotics model (vision-language-action) and cannot be run through the railwail API.

02

About Octo Base

TL;DRAs of September 23, 2026

Octo Base is a model by Other in the Robotics / VLA category. Octo Base is currently not available on Railwail.

Background

About UC Berkeley / Stanford (Octo Model Team)

Founded 2023 ยท Berkeley & Stanford, California, USA

The Octo project is a collaboration of academic labs led by Sergey Levine (UC Berkeley BAIR) and Chelsea Finn (Stanford IRIS), with contributions from CMU, Google DeepMind, and Toyota Research Institute. Octo was first released in May 2024 alongside the Open-X-Embodiment dataset effort, with the goal of producing a generalist, fully open-source robot policy that any researcher can fine-tune on a new robot in hours. Octo introduced the recipe of a transformer policy with a diffusion action head trained on 800k cross-embodiment demonstrations, and it has become a de-facto baseline in academic VLA / generalist-policy research. The team released both Octo-Small (27M) and Octo-Base (93M) under Apache-2.0, alongside code, checkpoints and a fine-tuning toolkit.

Visit UC Berkeley / Stanford (Octo Model Team)

Architecture

Transformer policy with diffusion action head (Vision-Language-Action)

Octo-Base is a transformer-based generalist robot policy. Inputs are tokenised RGB views and a natural-language instruction (encoded with a T5-base text encoder), interleaved with learnable readout tokens. The transformer trunk consumes this sequence and emits action latents that are decoded by a diffusion head producing continuous action chunks (default 4-step lookahead, 7-DoF end-effector deltas). The model was pretrained on roughly 800k demonstrations from 25 datasets in the Open-X-Embodiment collection, covering 9 robots, both single-arm and bimanual setups. Octo is intentionally embodiment-agnostic: action and proprioception spaces are encoded via shared adapters so the same backbone can be fine-tuned to new robots with as little as a few hundred demos. The diffusion head gives smooth, multimodal trajectories that outperform discrete-token VLAs on dexterous tasks at this scale.

Parameters
93M

Capabilities

  • Generalist VLA policy across many robot embodiments
  • Trained on ~800k demos from Open-X-Embodiment
  • Diffusion action head produces smooth continuous actions
  • Natural-language instruction conditioning (T5 encoder)
  • Multi-view image inputs (primary + wrist cameras)
  • Designed for fast fine-tuning on new robots and tasks
  • Apache-2.0 open weights, code and recipes
  • Strong academic baseline for VLA papers
  • Best for: research, fine-tuning to new embodiments, generalist-policy benchmarks.

Training & license

~800,000 robot trajectories drawn from 25 Open-X-Embodiment-compatible datasets across 9 robot embodiments (Franka, WidowX, Bridge, RT-1 Everyday Robots, Berkeley UR5, etc.). Trained on TPU v4 / v5 hardware.

License: Apache-2.0 - fully open weights, code, and dataset references. Research-friendly; commercial use permitted under the licence.

Safety testing: Octo is a research artifact; no formal red-teaming or RSP. Risks (unsafe robot motions) are mitigated at deployment time by application-specific safety controllers, not by the model itself.

Known limitations

  • Modest 93M scale - underperforms 7B+ VLAs on hard generalisation
  • Optimised for 7-DoF end-effector control - bimanual humanoid action spaces need adapters
  • Limited language reasoning relative to LLM-backed VLAs
  • Image resolution capped (256x256)
  • Trained mostly on Western lab data - geographic bias
  • Long-horizon planning requires external prompt decomposition
03

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

04

API

Call Octo Base with your Railwail API key. Use this model ID in the request:

Not available via the API

Robotics models run on robot hardware, not through the railwail API.

05

Specifications

Model ID
octo-base
Developer
Other
Input
Text, Image
Output
Robot actions
Model size
93M
License
Apache-2.0 - fully open weights, code, and dataset references. Research-friendly; commercial use permitted under the licence.
Catalog entry updated
September 23, 2026

Tags

  • berkeley
  • stanford
  • vla
  • robotics
  • research-only
  • open-weights
  • small
06

Use cases

What it is used for

  • Generalist policy fine-tuning for new robots
  • Academic research and ablations on VLAs
  • Baseline in cross-embodiment robot learning papers
  • Educational / coursework robotics projects
  • Bootstrapping data-efficient manipulation policies
  • Reproducible benchmarks on Open-X-Embodiment
07

Frequently asked questions

What is Octo Base?

Octo Base is a model by Other in the Robotics / VLA category. It is listed on Railwail but cannot be run at the moment.

How much does Octo Base cost on Railwail?

Octo Base cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

How fast is Octo Base?

There are not enough measured runs of Octo Base on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

When should I use Octo Base?

Octo Base belongs to the Robotics / VLA category. The category page lists the other models of this kind with their prices.

All models in Robotics / VLA

Can Octo Base process images?

Yes. Octo Base accepts images as input in addition to text.

Can I use Octo Base right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.