NVIDIA Cosmos-Predict-1

Robotică / VLAIndisponibil
de OtherID model: cosmos-predict-1

NVIDIA's world foundation model for physical AI. Diffusion-based video prediction for robotics simulation.

Status
Indisponibil
Intrare → ieșire
Text + Imagine → Acțiuni robot
Dezvoltator
Other
Actualizat
24 septembrie 2026

NVIDIA Cosmos-Predict-1 nu este disponibil în acest moment

Poți citi în continuare detaliile pe această pagină. Alege una dintre alternativele disponibile de mai jos pentru a rula imediat un model comparabil.

01

Playground

NVIDIA Cosmos-Predict-1

Model de cercetare

Indisponibil în prezent

NVIDIA Cosmos-Predict-1 este un model de robotică (vision-language-action) și nu poate fi rulat prin railwail API.

02

Despre NVIDIA Cosmos-Predict-1

Pe scurtDin 24 septembrie 2026

NVIDIA Cosmos-Predict-1 este un model de Other din categoria Robotică / VLA. NVIDIA Cosmos-Predict-1 nu este disponibil în prezent pe Railwail.

Fundal

Despre NVIDIA

Fondat 1993 · Santa Clara, California, USA

NVIDIA is the dominant supplier of GPUs for AI training and inference and runs a large in-house research organisation across robotics, simulation, and generative modelling. NVIDIA Cosmos was announced at CES 2025 as a family of 'World Foundation Models' (WFMs) for Physical AI - models that predict how the physical world evolves given video, language, and action conditioning. Cosmos is positioned as a developer platform for robotics and autonomous-vehicle teams to generate synthetic training data, run policy evaluations in simulation, and bootstrap Vision-Language-Action (VLA) pipelines. The 'Predict-1' track focuses on diffusion-based video-future prediction conditioned on text and/or first-frame inputs and ships in 7B and 14B parameter variants with open weights under the NVIDIA Open Model License.

Vizitează NVIDIA

Arhitectură

Diffusion-based world foundation model (text/video-to-video) for Physical AI

Cosmos-Predict-1 is a diffusion world model that predicts future video frames conditioned on text prompts, a starting frame, or short context clips. It uses a 3D causal video tokenizer (Cosmos Tokenizer) to compress video into spatio-temporal latents, then runs a Diffusion Transformer in latent space with cross-attention to text embeddings produced by a T5-XXL encoder. Training data is a curated corpus of ~20 million hours of driving, robotics, and human-activity video, filtered for motion quality, captioning coverage and safety. The model is not itself a VLA controller, but is the world-model backbone of NVIDIA's Cosmos stack: Cosmos-Predict generates rollouts; Cosmos-Reason adds VLM reasoning over predicted futures; and Cosmos-Transfer adapts simulation-to-real video. In a VLA pipeline it provides synthetic 'imagined' trajectories and dense reward / value signals, and is used to evaluate manipulation and driving policies offline at scale.

Parametri
7B and 14B variants (Predict-1)

Capabilități

  • Predicts future video conditioned on text, image, or video context
  • Two open-weight variants: Predict-1-7B and Predict-1-14B
  • Generates physically plausible motion for driving, manipulation, and humanoid scenes
  • Integrates with NVIDIA Isaac, Omniverse, and DRIVE pipelines
  • Used as synthetic data engine for VLA / autonomy training
  • Supports prompt upsampling via Cosmos-Reason VLM
  • Cosmos Tokenizer (3D causal VAE) can be reused as a video encoder
  • Released alongside Cosmos-Reason and Cosmos-Transfer for full Physical AI stack
  • Best for: synthetic data, world-model research, robotics simulation, AV training.

Antrenament & licență

Trained on ~20 million hours of curated physical-world video (driving, robotics manipulation, humanoid / first-person, navigation) sourced from licensed and open datasets, with multi-stage filtering for motion quality, caption alignment and safety. Text conditioning uses T5-XXL embeddings.

Licență: NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.

Teste de siguranță: NVIDIA applied dataset-level filters for unsafe content and faces, plus output filters and watermarking guidance. The model is positioned as a research artifact and not a deployed user-facing product, so no formal RSP-style policy is published.

Limitări cunoscute

  • Short rollouts (a few seconds) before drift dominates
  • Not a controller - cannot produce robot actions on its own
  • Heavy GPU footprint for the 14B variant
  • Domain skew toward driving / Western indoor scenes
  • Output is video, not joint commands - needs downstream policy
  • Restricted commercial license (Open Model License, not Apache)
03

Prețuri

Indisponibil în prezent. Nu există preț pentru acest model în acest moment, deci nu poate fi executat.

04

API

Apelează NVIDIA Cosmos-Predict-1 cu cheia ta API Railwail. Folosește acest ID de model în cerere:

Nu este disponibil prin API

Modelele de robotică rulează pe hardware de robot, nu prin API-ul railwail.

05

Specificații

ID model
cosmos-predict-1
Dezvoltator
Other
Intrare
Text, Imagine
Ieșire
Acțiuni robot
Dimensiune model
7B and 14B variants (Predict-1)
Licență
NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.
Intrare catalog actualizată
24 septembrie 2026

Etichete

  • nvidia
  • cosmos
  • vla
  • robotics
  • research-only
  • open-weights
  • world-model
06

Cazuri de utilizare

Pentru ce se folosește

  • Synthetic data generation for robotics and AV
  • World-model research and benchmarks
  • Offline policy evaluation in imagined rollouts
  • Simulation-to-real video adaptation pipelines
  • Pretraining backbone for VLA models
  • Academic embodied-AI research (research-only license)
07

Întrebări frecvente

Ce este NVIDIA Cosmos-Predict-1?

NVIDIA Cosmos-Predict-1 este un model de Other din categoria Robotică / VLA. Este listat pe Railwail, dar nu poate fi rulat în acest moment.

Cât costă NVIDIA Cosmos-Predict-1 pe Railwail?

NVIDIA Cosmos-Predict-1 nu poate fi rulat pe Railwail în acest moment, deci nu există preț curent. Alternativele disponibile cu prețuri sunt listate mai jos pe această pagină.

Cât de rapid este NVIDIA Cosmos-Predict-1?

Nu sunt suficiente rulări măsurate ale NVIDIA Cosmos-Predict-1 pe Railwail încă pentru a indica un timp de rulare. Depinde de intrare, de setări și de sarcina la furnizor.

Când ar trebui să folosesc NVIDIA Cosmos-Predict-1?

NVIDIA Cosmos-Predict-1 aparține categoriei Robotică / VLA. Pagina categoriei listează celelalte modele de acest fel cu prețurile lor.

Toate modelele: Robotică / VLA

Poate NVIDIA Cosmos-Predict-1 procesa imagini?

Da. NVIDIA Cosmos-Predict-1 acceptă imagini ca intrare, pe lângă text.

Pot folosi NVIDIA Cosmos-Predict-1 chiar acum?

Indisponibil în prezent. Pagina rămâne online; alternativele disponibile din aceeași categorie sunt listate mai jos.

Toate modelele printr-o singură API

O cheie API pentru fiecare model pe Railwail. Utilizarea se percepe din credite prepay, 1 credit = 0,01 USD.