NVIDIA Cosmos-Predict-1

Robotique / VLANon disponible
par OtherID du modèle: cosmos-predict-1

NVIDIA's world foundation model for physical AI. Diffusion-based video prediction for robotics simulation.

Statut
Non disponible
Entrée → Sortie
Texte + Image → Actions de robot
Développeur
Other
Mis à jour
24 septembre 2026

NVIDIA Cosmos-Predict-1 n'est actuellement pas disponible

Vous pouvez toujours consulter les détails sur cette page. Choisissez l'une des alternatives disponibles ci-dessous pour exécuter immédiatement un modèle comparable.

01

Playground

NVIDIA Cosmos-Predict-1

Modèle de recherche

Actuellement indisponible

NVIDIA Cosmos-Predict-1 est un modèle de robotique (vision-langage-action) et ne peut pas être exécuté via l'API railwail.

02

À propos de NVIDIA Cosmos-Predict-1

RésuméAu 24 septembre 2026

NVIDIA Cosmos-Predict-1 est un modèle de Other dans la catégorie Robotique / VLA. NVIDIA Cosmos-Predict-1 n'est actuellement pas disponible sur Railwail.

Arrière-plan

À propos de NVIDIA

Fondée 1993 · Santa Clara, California, USA

NVIDIA is the dominant supplier of GPUs for AI training and inference and runs a large in-house research organisation across robotics, simulation, and generative modelling. NVIDIA Cosmos was announced at CES 2025 as a family of 'World Foundation Models' (WFMs) for Physical AI - models that predict how the physical world evolves given video, language, and action conditioning. Cosmos is positioned as a developer platform for robotics and autonomous-vehicle teams to generate synthetic training data, run policy evaluations in simulation, and bootstrap Vision-Language-Action (VLA) pipelines. The 'Predict-1' track focuses on diffusion-based video-future prediction conditioned on text and/or first-frame inputs and ships in 7B and 14B parameter variants with open weights under the NVIDIA Open Model License.

Visiter NVIDIA

Architecture

Diffusion-based world foundation model (text/video-to-video) for Physical AI

Cosmos-Predict-1 is a diffusion world model that predicts future video frames conditioned on text prompts, a starting frame, or short context clips. It uses a 3D causal video tokenizer (Cosmos Tokenizer) to compress video into spatio-temporal latents, then runs a Diffusion Transformer in latent space with cross-attention to text embeddings produced by a T5-XXL encoder. Training data is a curated corpus of ~20 million hours of driving, robotics, and human-activity video, filtered for motion quality, captioning coverage and safety. The model is not itself a VLA controller, but is the world-model backbone of NVIDIA's Cosmos stack: Cosmos-Predict generates rollouts; Cosmos-Reason adds VLM reasoning over predicted futures; and Cosmos-Transfer adapts simulation-to-real video. In a VLA pipeline it provides synthetic 'imagined' trajectories and dense reward / value signals, and is used to evaluate manipulation and driving policies offline at scale.

Paramètres
7B and 14B variants (Predict-1)

Capacités

  • Predicts future video conditioned on text, image, or video context
  • Two open-weight variants: Predict-1-7B and Predict-1-14B
  • Generates physically plausible motion for driving, manipulation, and humanoid scenes
  • Integrates with NVIDIA Isaac, Omniverse, and DRIVE pipelines
  • Used as synthetic data engine for VLA / autonomy training
  • Supports prompt upsampling via Cosmos-Reason VLM
  • Cosmos Tokenizer (3D causal VAE) can be reused as a video encoder
  • Released alongside Cosmos-Reason and Cosmos-Transfer for full Physical AI stack
  • Best for: synthetic data, world-model research, robotics simulation, AV training.

Entraînement et licence

Trained on ~20 million hours of curated physical-world video (driving, robotics manipulation, humanoid / first-person, navigation) sourced from licensed and open datasets, with multi-stage filtering for motion quality, caption alignment and safety. Text conditioning uses T5-XXL embeddings.

Licence: NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.

Tests de sécurité: NVIDIA applied dataset-level filters for unsafe content and faces, plus output filters and watermarking guidance. The model is positioned as a research artifact and not a deployed user-facing product, so no formal RSP-style policy is published.

Limitations connues

  • Short rollouts (a few seconds) before drift dominates
  • Not a controller - cannot produce robot actions on its own
  • Heavy GPU footprint for the 14B variant
  • Domain skew toward driving / Western indoor scenes
  • Output is video, not joint commands - needs downstream policy
  • Restricted commercial license (Open Model License, not Apache)
03

Tarification

Actuellement indisponible. Il n'y a actuellement pas de prix pour ce modèle, il ne peut donc pas être exécuté.

04

API

Appelez NVIDIA Cosmos-Predict-1 avec votre clé API Railwail. Utilisez cet ID de modèle dans la requête :

Non disponible via l'API

Les modèles de robotique s'exécutent sur du matériel robotique, pas via l'API railwail.

05

Spécifications

ID du modèle
cosmos-predict-1
Développeur
Other
Catégorie
Robotique / VLA
Entrée
Texte, Image
Sortie
Actions de robot
Taille du modèle
7B and 14B variants (Predict-1)
Licence
NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.
Entrée du catalogue mise à jour
24 septembre 2026

Étiquettes

  • nvidia
  • cosmos
  • vla
  • robotics
  • research-only
  • open-weights
  • world-model
06

Cas d'usage

À quoi ça sert

  • Synthetic data generation for robotics and AV
  • World-model research and benchmarks
  • Offline policy evaluation in imagined rollouts
  • Simulation-to-real video adaptation pipelines
  • Pretraining backbone for VLA models
  • Academic embodied-AI research (research-only license)
07

Questions fréquemment posées

Qu'est-ce que NVIDIA Cosmos-Predict-1 ?

NVIDIA Cosmos-Predict-1 est un modèle de Other dans la catégorie Robotique / VLA. Il est répertorié sur Railwail mais ne peut pas être exécuté pour le moment.

Combien coûte NVIDIA Cosmos-Predict-1 sur Railwail ?

NVIDIA Cosmos-Predict-1 ne peut pas être exécuté sur Railwail pour le moment, il n'y a donc pas de prix actuel. Les alternatives disponibles avec leurs prix sont listées plus bas sur cette page.

Quelle est la vitesse de NVIDIA Cosmos-Predict-1 ?

Il n'y a pas encore assez d'exécutions mesurées de NVIDIA Cosmos-Predict-1 sur Railwail pour indiquer un temps d'exécution. Cela dépend de l'entrée, des paramètres et de la charge chez le fournisseur.

Quand utiliser NVIDIA Cosmos-Predict-1 ?

NVIDIA Cosmos-Predict-1 appartient à la catégorie Robotique / VLA. La page de catégorie liste les autres modèles de ce type avec leurs prix.

Tous les modèles : Robotique / VLA

NVIDIA Cosmos-Predict-1 peut-il traiter des images ?

Oui. NVIDIA Cosmos-Predict-1 accepte les images en entrée en plus du texte.

Puis-je utiliser NVIDIA Cosmos-Predict-1 maintenant ?

Actuellement indisponible. La page reste en ligne ; les alternatives disponibles de la même catégorie sont listées plus bas.

Tous les modèles via une seule API

Une clé API pour tous les modèles sur Railwail. L'utilisation est facturée à partir de crédits prépayés, 1 crédit = 0,01 $US.