NVIDIA Cosmos-Predict-1

Ρομποτική / VLAΜη διαθέσιμο
από OtherΑναγνωριστικό μοντέλου: cosmos-predict-1

NVIDIA's world foundation model for physical AI. Diffusion-based video prediction for robotics simulation.

Κατάσταση
Μη διαθέσιμο
Είσοδος → Έξοδος
Κείμενο + Εικόνα → Ενέργειες ρομπότ
Προγραμματιστής
Other
Ενημερώθηκε
24 Σεπτεμβρίου 2026

Το NVIDIA Cosmos-Predict-1 δεν είναι διαθέσιμο αυτή τη στιγμή

Μπορείτε να διαβάσετε τις λεπτομέρειες σε αυτήν τη σελίδα. Επιλέξτε μία από τις διαθέσιμες εναλλακτικές λύσεις παρακάτω για να εκτελέσετε αμέσως ένα συγκρίσιμο μοντέλο.

01

Playground

NVIDIA Cosmos-Predict-1

Μοντέλο έρευνας

Προς το παρόν μη διαθέσιμο

Το NVIDIA Cosmos-Predict-1 είναι ένα μοντέλο ρομποτικής (vision-language-action) και δεν μπορεί να εκτελεστεί μέσω του railwail API.

02

Σχετικά με το NVIDIA Cosmos-Predict-1

ΣύντομαΗμερομηνία: 24 Σεπτεμβρίου 2026

Το NVIDIA Cosmos-Predict-1 είναι ένα μοντέλο του Other στην κατηγορία Ρομποτική / VLA. Το NVIDIA Cosmos-Predict-1 δεν είναι διαθέσιμο στο Railwail αυτή τη στιγμή.

Φόντο

Σχετικά με NVIDIA

Ιδρύθηκε 1993 · Santa Clara, California, USA

NVIDIA is the dominant supplier of GPUs for AI training and inference and runs a large in-house research organisation across robotics, simulation, and generative modelling. NVIDIA Cosmos was announced at CES 2025 as a family of 'World Foundation Models' (WFMs) for Physical AI - models that predict how the physical world evolves given video, language, and action conditioning. Cosmos is positioned as a developer platform for robotics and autonomous-vehicle teams to generate synthetic training data, run policy evaluations in simulation, and bootstrap Vision-Language-Action (VLA) pipelines. The 'Predict-1' track focuses on diffusion-based video-future prediction conditioned on text and/or first-frame inputs and ships in 7B and 14B parameter variants with open weights under the NVIDIA Open Model License.

Επισκεφθείτε NVIDIA

Αρχιτεκτονική

Diffusion-based world foundation model (text/video-to-video) for Physical AI

Cosmos-Predict-1 is a diffusion world model that predicts future video frames conditioned on text prompts, a starting frame, or short context clips. It uses a 3D causal video tokenizer (Cosmos Tokenizer) to compress video into spatio-temporal latents, then runs a Diffusion Transformer in latent space with cross-attention to text embeddings produced by a T5-XXL encoder. Training data is a curated corpus of ~20 million hours of driving, robotics, and human-activity video, filtered for motion quality, captioning coverage and safety. The model is not itself a VLA controller, but is the world-model backbone of NVIDIA's Cosmos stack: Cosmos-Predict generates rollouts; Cosmos-Reason adds VLM reasoning over predicted futures; and Cosmos-Transfer adapts simulation-to-real video. In a VLA pipeline it provides synthetic 'imagined' trajectories and dense reward / value signals, and is used to evaluate manipulation and driving policies offline at scale.

Παράμετροι
7B and 14B variants (Predict-1)

Δυνατότητες

  • Predicts future video conditioned on text, image, or video context
  • Two open-weight variants: Predict-1-7B and Predict-1-14B
  • Generates physically plausible motion for driving, manipulation, and humanoid scenes
  • Integrates with NVIDIA Isaac, Omniverse, and DRIVE pipelines
  • Used as synthetic data engine for VLA / autonomy training
  • Supports prompt upsampling via Cosmos-Reason VLM
  • Cosmos Tokenizer (3D causal VAE) can be reused as a video encoder
  • Released alongside Cosmos-Reason and Cosmos-Transfer for full Physical AI stack
  • Best for: synthetic data, world-model research, robotics simulation, AV training.

Εκπαίδευση & άδεια

Trained on ~20 million hours of curated physical-world video (driving, robotics manipulation, humanoid / first-person, navigation) sourced from licensed and open datasets, with multi-stage filtering for motion quality, caption alignment and safety. Text conditioning uses T5-XXL embeddings.

Άδεια: NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.

Δοκιμές ασφάλειας: NVIDIA applied dataset-level filters for unsafe content and faces, plus output filters and watermarking guidance. The model is positioned as a research artifact and not a deployed user-facing product, so no formal RSP-style policy is published.

Γνωστοί περιορισμοί

  • Short rollouts (a few seconds) before drift dominates
  • Not a controller - cannot produce robot actions on its own
  • Heavy GPU footprint for the 14B variant
  • Domain skew toward driving / Western indoor scenes
  • Output is video, not joint commands - needs downstream policy
  • Restricted commercial license (Open Model License, not Apache)
03

Τιμολόγηση

Προς το παρόν μη διαθέσιμο. Δεν υπάρχει τιμή για αυτό το μοντέλο αυτή τη στιγμή, επομένως δεν μπορεί να εκτελεστεί.

04

API

Καλέστε το NVIDIA Cosmos-Predict-1 με το κλειδί Railwail API. Χρησιμοποιήστε αυτό το ID μοντέλου στο αίτημα:

Δεν είναι διαθέσιμο μέσω του API

Τα μοντέλα ρομποτικής εκτελούνται σε υλικό ρομπότ, όχι μέσω του railwail API.

05

Προδιαγραφές

ID μοντέλου
cosmos-predict-1
Ανάπτυξη
Other
Κατηγορία
Ρομποτική / VLA
Είσοδος
Κείμενο, Εικόνα
Έξοδος
Ενέργειες ρομπότ
Μέγεθος μοντέλου
7B and 14B variants (Predict-1)
Άδεια
NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.
Καταχώρηση καταλόγου ενημερώθηκε
24 Σεπτεμβρίου 2026

Ετικέτες

  • nvidia
  • cosmos
  • vla
  • robotics
  • research-only
  • open-weights
  • world-model
06

Περιπτώσεις χρήσης

Για τι χρησιμοποιείται

  • Synthetic data generation for robotics and AV
  • World-model research and benchmarks
  • Offline policy evaluation in imagined rollouts
  • Simulation-to-real video adaptation pipelines
  • Pretraining backbone for VLA models
  • Academic embodied-AI research (research-only license)
07

Συχνές ερωτήσεις

Τι είναι NVIDIA Cosmos-Predict-1;

Το NVIDIA Cosmos-Predict-1 είναι ένα μοντέλο του Other στην κατηγορία Ρομποτική / VLA. Είναι καταχωρημένο στο Railwail αλλά δεν μπορεί να εκτελεστεί αυτή τη στιγμή.

Πόσο κοστίζει το NVIDIA Cosmos-Predict-1 στο Railwail;

Το NVIDIA Cosmos-Predict-1 δεν μπορεί να εκτελεστεί στο Railwail αυτή τη στιγμή, επομένως δεν υπάρχει τρέχουσα τιμή. Διαθέσιμες εναλλακτικές λύσεις με τιμές παρατίθενται παρακάτω σε αυτήν τη σελίδα.

Πόσο γρήγορο είναι το NVIDIA Cosmos-Predict-1;

Δεν υπάρχουν αρκετές μετρημένες εκτελέσεις του NVIDIA Cosmos-Predict-1 στο Railwail ακόμα για να δηλωθεί ένας χρόνος εκτέλεσης. Εξαρτάται από την είσοδο, τις ρυθμίσεις και το φορτίο στον πάροχο.

Πότε πρέπει να χρησιμοποιήσω το NVIDIA Cosmos-Predict-1;

Το NVIDIA Cosmos-Predict-1 ανήκει στην κατηγορία Ρομποτική / VLA. Η σελίδα κατηγορίας παραθέτει τα άλλα μοντέλα αυτού του είδους με τις τιμές τους.

Όλα τα μοντέλα: Ρομποτική / VLA

Μπορεί το NVIDIA Cosmos-Predict-1 να επεξεργαστεί εικόνες;

Ναι. Το NVIDIA Cosmos-Predict-1 δέχεται εικόνες ως είσοδο εκτός από κείμενο.

Μπορώ να χρησιμοποιήσω το NVIDIA Cosmos-Predict-1 τώρα;

Προς το παρόν μη διαθέσιμο. Η σελίδα παραμένει ενεργή· διαθέσιμες εναλλακτικές λύσεις από την ίδια κατηγορία παρατίθενται παρακάτω.

Όλα τα μοντέλα μέσω ενός API

Ένα κλειδί API για κάθε μοντέλο στο Railwail. Η χρήση χρεώνεται από προπληρωμένα πιστωτικά, 1 πιστωτικό = 0,01 $.