Robotics / VLA

Vision-language-action models that turn camera images and instructions into robot actions.

Available
0of 15

15 models

  • UnavailableNew

    Action Chunking with Transformers, the imitation-learning policy from the ALOHA project, in the LeRobot implementation. It is trained from scratch on your own LeRobot dataset (no pretrained base model), predicts chunks of future joint actions from camera images and robot state, and trains quickly on a single GPU. Railwail offers it for fine-tuning only, not for hosted inference.

    actaloharobotics

    Modalities: Text, Image, Robot actions

    Deactivated

  • UnavailableNew

    NVIDIA's open 3B-parameter vision-language-action model for humanoids and robot arms: camera images, a language instruction and robot state in, continuous action vectors out. Weights under the NVIDIA Open Model License (commercial use allowed). Fine-tunable on your own LeRobot datasets.

    nvidiagrootvla

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Gemini Robotics (2025)

    Google DeepMind

    Unavailable

    Google DeepMind's vision-language-action model based on Gemini 2.0. Generalist robot policy with strong dexterity.

    geminivlarobotics

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Gemini Robotics-ER

    Google DeepMind

    Unavailable

    Embodied-reasoning variant of Gemini Robotics. Enhanced 3D spatial reasoning and trajectory planning.

    geminivlarobotics

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Google RT-2-X

    Google DeepMind

    Unavailable

    Google's VLA from RT-X collaboration. Trained on Open-X-Embodiment (22 robots, 527 skills), positive transfer.

    vlaroboticsresearch-only

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Unavailable

    HuggingFace's 450M VLA pretrained on 487 community LeRobot datasets. Runs on consumer GPUs.

    lerobotvlarobotics

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Unavailable

    NVIDIA's world foundation model for physical AI. Diffusion-based video prediction for robotics simulation.

    nvidiacosmosvla

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Unavailable

    NVIDIA's earlier 3B vision-language-action model (June 2025). Its weights are licensed for non-commercial use only, so Railwail does not run or fine-tune it. Use GR00T N1.7, which is commercially licensable.

    nvidiagrootvla

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Unavailable

    Berkeley/Stanford 93M transformer diffusion policy. Pretrained on 800k Open-X-Embodiment episodes.

    berkeleystanfordvla

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Unavailable

    Compact 27M variant of Octo. Faster inference on consumer GPUs, designed for low-latency control.

    berkeleyvlarobotics

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Unavailable

    Stanford/Berkeley open VLA trained on 970k Open-X-Embodiment episodes. Supports LoRA fine-tuning.

    stanfordberkeleyvla

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Physical Intelligence Pi-0-FAST

    Physical Intelligence

    Unavailable

    Autoregressive Ο€-0 variant using FAST action tokenizer. Faster inference at competitive task success.

    vlaroboticsresearch-only

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Physical Intelligence Ο€-0

    Physical Intelligence

    Unavailable

    Physical Intelligence's flagship VLA flow-matching policy. Generalist robot control, pretrained on 10k+ hrs robot data.

    vlaroboticsresearch-only

    Modalities: Text, Image, Robot actions

    Currently not offered

  • Physical Intelligence Ο€-0.5

    Physical Intelligence

    Unavailable

    Upgraded Ο€-0 with open-world generalization via knowledge insulation. Weights and fine-tuning open-sourced.

    vlaroboticsresearch-only

    Modalities: Text, Image, Robot actions

    Currently not offered

  • RDT-1B

    Other

    Unavailable

    Tsinghua's 1B diffusion-transformer bimanual manipulation policy. Predicts next 64 actions per inference.

    tsinghuavlarobotics

    Modalities: Text, Image, Robot actions

    Currently not offered

Vision-language-action models for robotics and embodied AI

Vision-language-action (VLA) models bridge perception, language, and motor control. A VLA takes camera frames plus a natural-language instruction ('pick up the red mug') and outputs low-level robot actions β€” joint angles, gripper commands, end-effector poses. Most are research artifacts from labs like Physical Intelligence, Google DeepMind, Stanford, and Berkeley.

Pricing, trade-offs and pitfalls

Pricing in this category is not yet standardized. Most of the models on this page run on dedicated GPU infrastructure β€” Vast.ai, Replicate, self-hosted β€” and you pay per second of inference compute rather than per call or per token; the cost per step depends on the GPU and the model size. On Railwail these models are listed for reference and cannot currently be run.

The trade-off triangle is generalization, latency, and physical scope. Larger VLAs (RT-2-X, OpenVLA-7B) generalize to novel objects and instructions but inference at 1-3 Hz, which is too slow for closed-loop dexterous control. Smaller distilled models (Octo, Ο€-0-fast, RDT-1B) hit 30-50 Hz but only generalize within their training distribution. For tabletop manipulation in a controlled cell, the small fast model is usually correct. For research that needs language and visual generalization, the larger model is.

Watch out for the sim-to-real gap: most VLA training data is collected in simulation or on specific robot embodiments. Deploying on a different arm, gripper, or camera geometry typically requires fine-tuning on a few hundred to a few thousand new demonstrations. Also watch out for safety β€” these models occasionally output unsafe joint trajectories; always run a low-level safety filter (joint limits, force limits, workspace bounds) between the policy and the hardware.

Top picks above cover the most generalizable research flagship, the cheapest run-on-shared-GPU option, the largest open-weights model, and the fastest realtime control policy. Commercial managed-API offerings will be added as providers launch them.

Typical tasks

  • Tabletop pick-and-place research
  • Mobile manipulation benchmarks
  • Bimanual coordination demos
  • Instruction-following teleop assistance
  • Sim-to-real policy transfer
  • Cross-embodiment benchmarking

Model comparisons

Frequently asked questions

Can I use these models commercially?

Most VLA models on this page are research-only β€” Apache 2.0 or MIT license on the code, restricted to non-commercial research on the weights. A few (Ο€-0-fast, RDT-1B) ship with broader licenses. Always read the model card before deploying on a paid product. Commercial managed-API offerings are expected to roll out across 2026.

What hardware do they run on?

Inference typically requires a single H100 or A100 GPU per robot at 10-50 Hz. Smaller distilled policies (Octo-small, Ο€-0-fast) can run on a single 4090 or A6000. For research, most labs run them on workstations adjacent to the robot. For production, expect to dedicate one GPU per active robot or one shared GPU across a small fleet.

How is inference billed?

On shared-GPU platforms (Vast.ai, Replicate) you pay per second of compute; the cost per inference step depends on the GPU and the model size. Self-hosted on your own hardware is electricity plus depreciation. On Railwail, the VLA models are not currently available to run.

What robot embodiments are supported?

Most VLAs are trained on specific platforms β€” Franka Panda, UR5, ALOHA, mobile ALOHA, Cobot Magic, etc. Cross-embodiment generalization is improving (Octo and RT-X were explicit attempts) but deploying on a new arm still typically requires 100-1,000 fine-tuning demonstrations. Check the model card for trained embodiments.

Can they handle dexterous manipulation?

Tabletop pick-and-place is reliable on most VLAs. Multi-finger dexterity, in-hand manipulation, and tool use are still hard β€” they work in demos but generalize poorly. Ο€-0 and RT-2 show the strongest dexterity to date in open research; expect rapid progress through 2026.

What's the difference between VLA and a regular policy network?

A regular policy maps observations to actions. A VLA additionally conditions on a natural-language instruction, so the same policy can do 'pick up the red mug' and 'pick up the blue cup' from the same model. This shifts complexity from per-task training to large-scale instruction-action pretraining.

How do I fine-tune for my robot?

Collect 100-1,000 teleoperated demonstrations of your target tasks, then run supervised fine-tuning (typically LoRA) on the pretrained checkpoint. Most repositories include a fine-tuning script. Plan for 4-24 hours of GPU time per fine-tune on a single H100, plus a few days of evaluation iteration.

What does the future look like for commercial VLAs?

Physical Intelligence, Skild AI, Covariant, and a handful of stealth labs are explicitly building general-purpose commercial VLAs with managed APIs. Expect the first commercial offerings (likely vertically integrated with specific robot OEMs) to ship across 2026 and 2027. Railwail will list them here as they launch.

Build with one API

Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.