RDT-1B

Robotics / VLAUnavailable
by OtherModel ID: rdt-1b

Tsinghua's 1B diffusion-transformer bimanual manipulation policy. Predicts next 64 actions per inference.

Status
Unavailable
Input โ†’ output
Text + Image โ†’ Robot actions
Developer
Other
Updated
September 23, 2026

RDT-1B is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

01

Playground

RDT-1B

Research model

Currently unavailable

RDT-1B is a robotics model (vision-language-action) and cannot be run through the railwail API.

02

About RDT-1B

TL;DRAs of September 23, 2026

RDT-1B is a model by Other in the Robotics / VLA category. RDT-1B is currently not available on Railwail.

Background

About Tsinghua University (TSAIL / IIIS)

Founded 1911 ยท Beijing, China

Robotics Diffusion Transformer (RDT) is a generalist bimanual manipulation policy developed at Tsinghua University's TSAIL / Institute for Interdisciplinary Information Sciences (IIIS), home of Jun Zhu's diffusion-modelling group. RDT-1B, introduced in October 2024, is one of the first publicly released billion-scale diffusion-based Vision-Language-Action models, specifically designed for two-arm robots such as Aloha, Mobile Aloha and a custom bimanual platform used by the authors. The project is positioned as a Chinese academic counterpart to ฯ€-0 and OpenVLA, with open weights released on Hugging Face under a permissive licence and the explicit aim of enabling fully reproducible bimanual VLA research.

Visit Tsinghua University (TSAIL / IIIS)

Architecture

Diffusion Transformer Vision-Language-Action policy for bimanual manipulation

RDT-1B is a 1-billion-parameter Diffusion Transformer (DiT) trained as a Vision-Language-Action policy. Inputs are multi-view RGB observations (left + right + overhead), proprioception for both arms and any gripper / mobile-base degrees of freedom, plus a natural-language instruction encoded by a text encoder. The conditioning tokens are fed through a transformer trunk, while a diffusion head denoises continuous action chunks for both arms in a unified action space, allowing dual-arm coordinated motion. Pretraining is done in two stages: a large multi-robot pretraining phase on >1M episodes drawn from public datasets including Open-X-Embodiment and curated bimanual corpora, followed by fine-tuning on the authors' own 6,000-episode bimanual dataset spanning ~300 tasks. RDT-1B reports strong results on dexterous bimanual tasks such as folding T-shirts, pouring, and tool use.

Parameters
1B

Capabilities

  • 1B-parameter Diffusion Transformer VLA
  • Designed for bimanual manipulation (Aloha-class robots)
  • Trained on >1M cross-embodiment episodes + 6k bimanual demos
  • Continuous action chunks for both arms in a unified space
  • Diffusion head produces smooth coordinated motion
  • Open weights on Hugging Face (permissive licence)
  • Strong results on folding, pouring and tool use
  • Reproducible training and evaluation code
  • Best for: bimanual manipulation research, two-arm fine-tuning.

Training & license

Pretraining on >1 million robot episodes from Open-X-Embodiment and other public datasets, followed by fine-tuning on a curated bimanual dataset of ~6,000 episodes covering ~300 tasks collected with Aloha-class hardware.

License: Open weights released on Hugging Face under a permissive (CC-BY-NC-style) licence; primarily intended for research use.

Safety testing: Academic research artifact; no formal red-teaming or RSP. Safety is the responsibility of downstream deployers (Aloha hardware already includes torque limits and e-stop).

Known limitations

  • Primarily targets bimanual Aloha-class hardware
  • Requires diffusion sampling at inference (multiple steps)
  • Limited language reasoning compared to LLM-backed VLAs
  • Generalisation to single-arm or mobile platforms needs adapters
  • Mostly indoor-lab evaluation
  • Smaller pretraining text corpus than RT-2-X / OpenVLA
03

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

04

API

Call RDT-1B with your Railwail API key. Use this model ID in the request:

Not available via the API

Robotics models run on robot hardware, not through the railwail API.

05

Specifications

Model ID
rdt-1b
Developer
Other
Input
Text, Image
Output
Robot actions
Model size
1B
License
Open weights released on Hugging Face under a permissive (CC-BY-NC-style) licence; primarily intended for research use.
Catalog entry updated
September 23, 2026

Tags

  • tsinghua
  • vla
  • robotics
  • bimanual
  • research-only
  • open-weights
  • diffusion
06

Use cases

What it is used for

  • Bimanual manipulation research
  • Aloha / Mobile Aloha policy fine-tuning
  • Diffusion-policy ablations and benchmarks
  • Open-source baseline for two-arm VLAs
  • Coordinated dual-arm task learning
  • Academic studies of large diffusion policies
07

Frequently asked questions

What is RDT-1B?

RDT-1B is a model by Other in the Robotics / VLA category. It is listed on Railwail but cannot be run at the moment.

How much does RDT-1B cost on Railwail?

RDT-1B cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

How fast is RDT-1B?

There are not enough measured runs of RDT-1B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

When should I use RDT-1B?

RDT-1B belongs to the Robotics / VLA category. The category page lists the other models of this kind with their prices.

All models in Robotics / VLA

Can RDT-1B process images?

Yes. RDT-1B accepts images as input in addition to text.

Can I use RDT-1B right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.