NVIDIA Cosmos-Predict-1

Robotik / VLANicht verfügbar
von OtherModell-ID: cosmos-predict-1

NVIDIA's world foundation model for physical AI. Diffusion-based video prediction for robotics simulation.

Status
Nicht verfügbar
Eingabe → Ausgabe
Text + Bild → Roboteraktionen
Entwickler
Other
Aktualisiert
24. September 2026

NVIDIA Cosmos-Predict-1 ist derzeit nicht verfügbar

Die Angaben auf dieser Seite kannst du weiter nachlesen. Mit einer der verfügbaren Alternativen unten kannst du sofort ein vergleichbares Modell nutzen.

01

Playground

NVIDIA Cosmos-Predict-1

Forschungsmodell

Derzeit nicht verfügbar

NVIDIA Cosmos-Predict-1 ist ein Robotik-Modell (Vision-Language-Action) und lässt sich nicht über die railwail-API ausführen.

02

Über NVIDIA Cosmos-Predict-1

Kurz gesagtStand: 24. September 2026

NVIDIA Cosmos-Predict-1 ist ein Modell von Other aus der Kategorie Robotik / VLA. Über Railwail ist NVIDIA Cosmos-Predict-1 derzeit nicht verfügbar.

Hintergrund

Über NVIDIA

Gegründet 1993 · Santa Clara, California, USA

NVIDIA is the dominant supplier of GPUs for AI training and inference and runs a large in-house research organisation across robotics, simulation, and generative modelling. NVIDIA Cosmos was announced at CES 2025 as a family of 'World Foundation Models' (WFMs) for Physical AI - models that predict how the physical world evolves given video, language, and action conditioning. Cosmos is positioned as a developer platform for robotics and autonomous-vehicle teams to generate synthetic training data, run policy evaluations in simulation, and bootstrap Vision-Language-Action (VLA) pipelines. The 'Predict-1' track focuses on diffusion-based video-future prediction conditioned on text and/or first-frame inputs and ships in 7B and 14B parameter variants with open weights under the NVIDIA Open Model License.

NVIDIA besuchen

Architektur

Diffusion-based world foundation model (text/video-to-video) for Physical AI

Cosmos-Predict-1 is a diffusion world model that predicts future video frames conditioned on text prompts, a starting frame, or short context clips. It uses a 3D causal video tokenizer (Cosmos Tokenizer) to compress video into spatio-temporal latents, then runs a Diffusion Transformer in latent space with cross-attention to text embeddings produced by a T5-XXL encoder. Training data is a curated corpus of ~20 million hours of driving, robotics, and human-activity video, filtered for motion quality, captioning coverage and safety. The model is not itself a VLA controller, but is the world-model backbone of NVIDIA's Cosmos stack: Cosmos-Predict generates rollouts; Cosmos-Reason adds VLM reasoning over predicted futures; and Cosmos-Transfer adapts simulation-to-real video. In a VLA pipeline it provides synthetic 'imagined' trajectories and dense reward / value signals, and is used to evaluate manipulation and driving policies offline at scale.

Parameter
7B and 14B variants (Predict-1)

Funktionen

  • Predicts future video conditioned on text, image, or video context
  • Two open-weight variants: Predict-1-7B and Predict-1-14B
  • Generates physically plausible motion for driving, manipulation, and humanoid scenes
  • Integrates with NVIDIA Isaac, Omniverse, and DRIVE pipelines
  • Used as synthetic data engine for VLA / autonomy training
  • Supports prompt upsampling via Cosmos-Reason VLM
  • Cosmos Tokenizer (3D causal VAE) can be reused as a video encoder
  • Released alongside Cosmos-Reason and Cosmos-Transfer for full Physical AI stack
  • Best for: synthetic data, world-model research, robotics simulation, AV training.

Training & Lizenz

Trained on ~20 million hours of curated physical-world video (driving, robotics manipulation, humanoid / first-person, navigation) sourced from licensed and open datasets, with multi-stage filtering for motion quality, caption alignment and safety. Text conditioning uses T5-XXL embeddings.

Lizenz: NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.

Sicherheitstests: NVIDIA applied dataset-level filters for unsafe content and faces, plus output filters and watermarking guidance. The model is positioned as a research artifact and not a deployed user-facing product, so no formal RSP-style policy is published.

Bekannte Einschränkungen

  • Short rollouts (a few seconds) before drift dominates
  • Not a controller - cannot produce robot actions on its own
  • Heavy GPU footprint for the 14B variant
  • Domain skew toward driving / Western indoor scenes
  • Output is video, not joint commands - needs downstream policy
  • Restricted commercial license (Open Model License, not Apache)
03

Preise

Derzeit nicht verfügbar. Für dieses Modell gibt es derzeit keinen Preis, deshalb lässt es sich nicht ausführen.

04

API

Rufe NVIDIA Cosmos-Predict-1 mit deinem Railwail-API-Schlüssel auf. Diese Modell-ID gehört in die Anfrage:

Nicht über die API verfügbar

Robotik-Modelle laufen auf Roboter-Hardware, nicht über die railwail-API.

05

Spezifikationen

Modell-ID
cosmos-predict-1
Entwickler
Other
Kategorie
Robotik / VLA
Eingabe
Text, Bild
Ausgabe
Roboteraktionen
Modellgröße
7B and 14B variants (Predict-1)
Lizenz
NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.
Katalogeintrag aktualisiert
24. September 2026

Schlagwörter

  • nvidia
  • cosmos
  • vla
  • robotics
  • research-only
  • open-weights
  • world-model
06

Einsatzgebiete

Wofür es genutzt wird

  • Synthetic data generation for robotics and AV
  • World-model research and benchmarks
  • Offline policy evaluation in imagined rollouts
  • Simulation-to-real video adaptation pipelines
  • Pretraining backbone for VLA models
  • Academic embodied-AI research (research-only license)
07

Häufige Fragen

Was ist NVIDIA Cosmos-Predict-1?

NVIDIA Cosmos-Predict-1 ist ein Modell von Other aus der Kategorie Robotik / VLA. Es steht im Railwail-Katalog, lässt sich derzeit aber nicht ausführen.

Was kostet NVIDIA Cosmos-Predict-1 bei Railwail?

NVIDIA Cosmos-Predict-1 lässt sich über Railwail derzeit nicht ausführen, deshalb gibt es keinen aktuellen Preis. Verfügbare Alternativen mit Preisen stehen weiter unten auf dieser Seite.

Wie schnell ist NVIDIA Cosmos-Predict-1?

Für NVIDIA Cosmos-Predict-1 gibt es bei Railwail noch zu wenige gemessene Läufe, um eine Laufzeit anzugeben. Sie hängt von der Eingabe, den Einstellungen und der Auslastung beim Anbieter ab.

Wann sollte ich NVIDIA Cosmos-Predict-1 nutzen?

NVIDIA Cosmos-Predict-1 gehört zur Kategorie Robotik / VLA. Die Kategorieseite listet die anderen Modelle dieser Art mit ihren Preisen.

Alle Modelle: Robotik / VLA

Kann NVIDIA Cosmos-Predict-1 Bilder verarbeiten?

Ja. NVIDIA Cosmos-Predict-1 nimmt neben Text auch Bilder als Eingabe an.

Kann ich NVIDIA Cosmos-Predict-1 gerade nutzen?

Derzeit nicht verfügbar. Die Seite bleibt online; verfügbare Alternativen aus derselben Kategorie stehen weiter unten.

Alle Modelle über eine API

Ein API-Schlüssel für alle Modelle auf Railwail. Abgerechnet wird über vorab gekaufte Credits, 1 Credit = $ 0,01.