NVIDIA Cosmos-Predict-1

Robotica / VLANon disponibile
di OtherID modello: cosmos-predict-1

NVIDIA's world foundation model for physical AI. Diffusion-based video prediction for robotics simulation.

Stato
Non disponibile
Input → output
Testo + Immagine → Azioni robot
Sviluppatore
Other
Aggiornato
24 settembre 2026

NVIDIA Cosmos-Predict-1 non è attualmente disponibile

Puoi comunque leggere i dettagli su questa pagina. Scegli una delle alternative disponibili di seguito per eseguire subito un modello comparabile.

01

Playground

NVIDIA Cosmos-Predict-1

Modello di ricerca

Attualmente non disponibile

NVIDIA Cosmos-Predict-1 è un modello di robotica (vision-language-action) e non può essere eseguito tramite l'API railwail.

02

Informazioni su NVIDIA Cosmos-Predict-1

RiassuntoA partire da 24 settembre 2026

NVIDIA Cosmos-Predict-1 è un modello di Other nella categoria Robotica / VLA. NVIDIA Cosmos-Predict-1 non è attualmente disponibile su Railwail.

Sfondo

Informazioni su NVIDIA

Fondato 1993 · Santa Clara, California, USA

NVIDIA is the dominant supplier of GPUs for AI training and inference and runs a large in-house research organisation across robotics, simulation, and generative modelling. NVIDIA Cosmos was announced at CES 2025 as a family of 'World Foundation Models' (WFMs) for Physical AI - models that predict how the physical world evolves given video, language, and action conditioning. Cosmos is positioned as a developer platform for robotics and autonomous-vehicle teams to generate synthetic training data, run policy evaluations in simulation, and bootstrap Vision-Language-Action (VLA) pipelines. The 'Predict-1' track focuses on diffusion-based video-future prediction conditioned on text and/or first-frame inputs and ships in 7B and 14B parameter variants with open weights under the NVIDIA Open Model License.

Visita NVIDIA

Architettura

Diffusion-based world foundation model (text/video-to-video) for Physical AI

Cosmos-Predict-1 is a diffusion world model that predicts future video frames conditioned on text prompts, a starting frame, or short context clips. It uses a 3D causal video tokenizer (Cosmos Tokenizer) to compress video into spatio-temporal latents, then runs a Diffusion Transformer in latent space with cross-attention to text embeddings produced by a T5-XXL encoder. Training data is a curated corpus of ~20 million hours of driving, robotics, and human-activity video, filtered for motion quality, captioning coverage and safety. The model is not itself a VLA controller, but is the world-model backbone of NVIDIA's Cosmos stack: Cosmos-Predict generates rollouts; Cosmos-Reason adds VLM reasoning over predicted futures; and Cosmos-Transfer adapts simulation-to-real video. In a VLA pipeline it provides synthetic 'imagined' trajectories and dense reward / value signals, and is used to evaluate manipulation and driving policies offline at scale.

Parametri
7B and 14B variants (Predict-1)

Capacità

  • Predicts future video conditioned on text, image, or video context
  • Two open-weight variants: Predict-1-7B and Predict-1-14B
  • Generates physically plausible motion for driving, manipulation, and humanoid scenes
  • Integrates with NVIDIA Isaac, Omniverse, and DRIVE pipelines
  • Used as synthetic data engine for VLA / autonomy training
  • Supports prompt upsampling via Cosmos-Reason VLM
  • Cosmos Tokenizer (3D causal VAE) can be reused as a video encoder
  • Released alongside Cosmos-Reason and Cosmos-Transfer for full Physical AI stack
  • Best for: synthetic data, world-model research, robotics simulation, AV training.

Addestramento e licenza

Trained on ~20 million hours of curated physical-world video (driving, robotics manipulation, humanoid / first-person, navigation) sourced from licensed and open datasets, with multi-stage filtering for motion quality, caption alignment and safety. Text conditioning uses T5-XXL embeddings.

Licenza: NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.

Test di sicurezza: NVIDIA applied dataset-level filters for unsafe content and faces, plus output filters and watermarking guidance. The model is positioned as a research artifact and not a deployed user-facing product, so no formal RSP-style policy is published.

Limitazioni note

  • Short rollouts (a few seconds) before drift dominates
  • Not a controller - cannot produce robot actions on its own
  • Heavy GPU footprint for the 14B variant
  • Domain skew toward driving / Western indoor scenes
  • Output is video, not joint commands - needs downstream policy
  • Restricted commercial license (Open Model License, not Apache)
03

Prezzi

Attualmente non disponibile. Al momento non c'è un prezzo per questo modello, quindi non può essere eseguito.

04

API

Chiama NVIDIA Cosmos-Predict-1 con la tua chiave API Railwail. Usa questo ID modello nella richiesta:

Non disponibile tramite l'API

I modelli di robotica vengono eseguiti su hardware robotico, non tramite l'API railwail.

05

Specifiche

ID modello
cosmos-predict-1
Sviluppatore
Other
Input
Testo, Immagine
Output
Azioni robot
Dimensione del modello
7B and 14B variants (Predict-1)
Licenza
NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.
Voce di catalogo aggiornata
24 settembre 2026

Etichette

  • nvidia
  • cosmos
  • vla
  • robotics
  • research-only
  • open-weights
  • world-model
06

Casi d'uso

A cosa serve

  • Synthetic data generation for robotics and AV
  • World-model research and benchmarks
  • Offline policy evaluation in imagined rollouts
  • Simulation-to-real video adaptation pipelines
  • Pretraining backbone for VLA models
  • Academic embodied-AI research (research-only license)
07

Domande frequenti

Cos'è NVIDIA Cosmos-Predict-1?

NVIDIA Cosmos-Predict-1 è un modello di Other nella categoria Robotica / VLA. È elencato su Railwail ma non può essere eseguito al momento.

Quanto costa NVIDIA Cosmos-Predict-1 su Railwail?

NVIDIA Cosmos-Predict-1 non può essere eseguito su Railwail al momento, quindi non c'è un prezzo attuale. Le alternative disponibili con i prezzi sono elencate più in basso in questa pagina.

Quanto è veloce NVIDIA Cosmos-Predict-1?

Non ci sono ancora abbastanza esecuzioni misurate di NVIDIA Cosmos-Predict-1 su Railwail per indicare un tempo di esecuzione. Dipende dall'input, dalle impostazioni e dal carico presso il provider.

Quando devo usare NVIDIA Cosmos-Predict-1?

NVIDIA Cosmos-Predict-1 appartiene alla categoria Robotica / VLA. La pagina della categoria elenca gli altri modelli di questo tipo con i loro prezzi.

Tutti i modelli: Robotica / VLA

NVIDIA Cosmos-Predict-1 può elaborare immagini?

Sì. NVIDIA Cosmos-Predict-1 accetta immagini come input oltre al testo.

Posso usare NVIDIA Cosmos-Predict-1 adesso?

Attualmente non disponibile. La pagina rimane online; le alternative disponibili della stessa categoria sono elencate più in basso.

Tutti i modelli tramite un'API

Una chiave API per tutti i modelli su Railwail. L'utilizzo viene addebitato da crediti prepagati, 1 credito = 0,01 USD.