NVIDIA Cosmos-Predict-1

Robotics / VLAUnavailable
by OtherModel ID: cosmos-predict-1

NVIDIA's world foundation model for physical AI. Diffusion-based video prediction for robotics simulation.

Status
Unavailable
Input β†’ output
Text + Image β†’ Robot actions
Developer
Other
Updated
September 24, 2026

NVIDIA Cosmos-Predict-1 is currently unavailable

You can still read the details on this page. Pick one of the available alternatives below to run a comparable model right away.

01

Playground

NVIDIA Cosmos-Predict-1

Research model

Currently unavailable

NVIDIA Cosmos-Predict-1 is a robotics model (vision-language-action) and cannot be run through the railwail API.

02

About NVIDIA Cosmos-Predict-1

TL;DRAs of September 24, 2026

NVIDIA Cosmos-Predict-1 is a model by Other in the Robotics / VLA category. NVIDIA Cosmos-Predict-1 is currently not available on Railwail.

Background

About NVIDIA

Founded 1993 Β· Santa Clara, California, USA

NVIDIA is the dominant supplier of GPUs for AI training and inference and runs a large in-house research organisation across robotics, simulation, and generative modelling. NVIDIA Cosmos was announced at CES 2025 as a family of 'World Foundation Models' (WFMs) for Physical AI - models that predict how the physical world evolves given video, language, and action conditioning. Cosmos is positioned as a developer platform for robotics and autonomous-vehicle teams to generate synthetic training data, run policy evaluations in simulation, and bootstrap Vision-Language-Action (VLA) pipelines. The 'Predict-1' track focuses on diffusion-based video-future prediction conditioned on text and/or first-frame inputs and ships in 7B and 14B parameter variants with open weights under the NVIDIA Open Model License.

Visit NVIDIA

Architecture

Diffusion-based world foundation model (text/video-to-video) for Physical AI

Cosmos-Predict-1 is a diffusion world model that predicts future video frames conditioned on text prompts, a starting frame, or short context clips. It uses a 3D causal video tokenizer (Cosmos Tokenizer) to compress video into spatio-temporal latents, then runs a Diffusion Transformer in latent space with cross-attention to text embeddings produced by a T5-XXL encoder. Training data is a curated corpus of ~20 million hours of driving, robotics, and human-activity video, filtered for motion quality, captioning coverage and safety. The model is not itself a VLA controller, but is the world-model backbone of NVIDIA's Cosmos stack: Cosmos-Predict generates rollouts; Cosmos-Reason adds VLM reasoning over predicted futures; and Cosmos-Transfer adapts simulation-to-real video. In a VLA pipeline it provides synthetic 'imagined' trajectories and dense reward / value signals, and is used to evaluate manipulation and driving policies offline at scale.

Parameters
7B and 14B variants (Predict-1)

Capabilities

  • Predicts future video conditioned on text, image, or video context
  • Two open-weight variants: Predict-1-7B and Predict-1-14B
  • Generates physically plausible motion for driving, manipulation, and humanoid scenes
  • Integrates with NVIDIA Isaac, Omniverse, and DRIVE pipelines
  • Used as synthetic data engine for VLA / autonomy training
  • Supports prompt upsampling via Cosmos-Reason VLM
  • Cosmos Tokenizer (3D causal VAE) can be reused as a video encoder
  • Released alongside Cosmos-Reason and Cosmos-Transfer for full Physical AI stack
  • Best for: synthetic data, world-model research, robotics simulation, AV training.

Training & license

Trained on ~20 million hours of curated physical-world video (driving, robotics manipulation, humanoid / first-person, navigation) sourced from licensed and open datasets, with multi-stage filtering for motion quality, caption alignment and safety. Text conditioning uses T5-XXL embeddings.

License: NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.

Safety testing: NVIDIA applied dataset-level filters for unsafe content and faces, plus output filters and watermarking guidance. The model is positioned as a research artifact and not a deployed user-facing product, so no formal RSP-style policy is published.

Known limitations

  • Short rollouts (a few seconds) before drift dominates
  • Not a controller - cannot produce robot actions on its own
  • Heavy GPU footprint for the 14B variant
  • Domain skew toward driving / Western indoor scenes
  • Output is video, not joint commands - needs downstream policy
  • Restricted commercial license (Open Model License, not Apache)
03

Pricing

Currently unavailable. There is no price for this model at the moment, so it cannot be run.

04

API

Call NVIDIA Cosmos-Predict-1 with your Railwail API key. Use this model ID in the request:

Not available via the API

Robotics models run on robot hardware, not through the railwail API.

05

Specifications

Model ID
cosmos-predict-1
Developer
Other
Input
Text, Image
Output
Robot actions
Model size
7B and 14B variants (Predict-1)
License
NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.
Catalog entry updated
September 24, 2026

Tags

  • nvidia
  • cosmos
  • vla
  • robotics
  • research-only
  • open-weights
  • world-model
06

Use cases

What it is used for

  • Synthetic data generation for robotics and AV
  • World-model research and benchmarks
  • Offline policy evaluation in imagined rollouts
  • Simulation-to-real video adaptation pipelines
  • Pretraining backbone for VLA models
  • Academic embodied-AI research (research-only license)
07

Frequently asked questions

What is NVIDIA Cosmos-Predict-1?

NVIDIA Cosmos-Predict-1 is a model by Other in the Robotics / VLA category. It is listed on Railwail but cannot be run at the moment.

How much does NVIDIA Cosmos-Predict-1 cost on Railwail?

NVIDIA Cosmos-Predict-1 cannot be run on Railwail at the moment, so there is no current price. Available alternatives with prices are listed further down this page.

How fast is NVIDIA Cosmos-Predict-1?

There are not enough measured runs of NVIDIA Cosmos-Predict-1 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

When should I use NVIDIA Cosmos-Predict-1?

NVIDIA Cosmos-Predict-1 belongs to the Robotics / VLA category. The category page lists the other models of this kind with their prices.

All models in Robotics / VLA

Can NVIDIA Cosmos-Predict-1 process images?

Yes. NVIDIA Cosmos-Predict-1 accepts images as input in addition to text.

Can I use NVIDIA Cosmos-Predict-1 right now?

Currently unavailable. The page stays online; available alternatives from the same category are listed further down.

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.