NVIDIA Cosmos-Predict-1

Robotyka / VLANiedostępne
od OtherID modelu: cosmos-predict-1

NVIDIA's world foundation model for physical AI. Diffusion-based video prediction for robotics simulation.

Status
Niedostępne
Wejście → Wyjście
Tekst + Obraz → Akcje robota
Deweloper
Other
Zaktualizowano
24 września 2026

NVIDIA Cosmos-Predict-1 jest obecnie niedostępny

Możesz nadal przeczytać szczegóły na tej stronie. Wybierz jedną z dostępnych alternatyw poniżej, aby od razu uruchomić porównywalny model.

01

Playground

NVIDIA Cosmos-Predict-1

Model badawczy

Niedostępny

NVIDIA Cosmos-Predict-1 to model robotyki (vision-language-action) i nie można go uruchomić przez API railwail.

02

O NVIDIA Cosmos-Predict-1

Krótko mówiącStan na 24 września 2026

NVIDIA Cosmos-Predict-1 to model opracowany przez Other w kategorii Robotyka / VLA. NVIDIA Cosmos-Predict-1 nie jest obecnie dostępny w serwisie Railwail.

Tło

O NVIDIA

Założona 1993 · Santa Clara, California, USA

NVIDIA is the dominant supplier of GPUs for AI training and inference and runs a large in-house research organisation across robotics, simulation, and generative modelling. NVIDIA Cosmos was announced at CES 2025 as a family of 'World Foundation Models' (WFMs) for Physical AI - models that predict how the physical world evolves given video, language, and action conditioning. Cosmos is positioned as a developer platform for robotics and autonomous-vehicle teams to generate synthetic training data, run policy evaluations in simulation, and bootstrap Vision-Language-Action (VLA) pipelines. The 'Predict-1' track focuses on diffusion-based video-future prediction conditioned on text and/or first-frame inputs and ships in 7B and 14B parameter variants with open weights under the NVIDIA Open Model License.

Odwiedź NVIDIA

Architektura

Diffusion-based world foundation model (text/video-to-video) for Physical AI

Cosmos-Predict-1 is a diffusion world model that predicts future video frames conditioned on text prompts, a starting frame, or short context clips. It uses a 3D causal video tokenizer (Cosmos Tokenizer) to compress video into spatio-temporal latents, then runs a Diffusion Transformer in latent space with cross-attention to text embeddings produced by a T5-XXL encoder. Training data is a curated corpus of ~20 million hours of driving, robotics, and human-activity video, filtered for motion quality, captioning coverage and safety. The model is not itself a VLA controller, but is the world-model backbone of NVIDIA's Cosmos stack: Cosmos-Predict generates rollouts; Cosmos-Reason adds VLM reasoning over predicted futures; and Cosmos-Transfer adapts simulation-to-real video. In a VLA pipeline it provides synthetic 'imagined' trajectories and dense reward / value signals, and is used to evaluate manipulation and driving policies offline at scale.

Parametry
7B and 14B variants (Predict-1)

Możliwości

  • Predicts future video conditioned on text, image, or video context
  • Two open-weight variants: Predict-1-7B and Predict-1-14B
  • Generates physically plausible motion for driving, manipulation, and humanoid scenes
  • Integrates with NVIDIA Isaac, Omniverse, and DRIVE pipelines
  • Used as synthetic data engine for VLA / autonomy training
  • Supports prompt upsampling via Cosmos-Reason VLM
  • Cosmos Tokenizer (3D causal VAE) can be reused as a video encoder
  • Released alongside Cosmos-Reason and Cosmos-Transfer for full Physical AI stack
  • Best for: synthetic data, world-model research, robotics simulation, AV training.

Trening i licencja

Trained on ~20 million hours of curated physical-world video (driving, robotics manipulation, humanoid / first-person, navigation) sourced from licensed and open datasets, with multi-stage filtering for motion quality, caption alignment and safety. Text conditioning uses T5-XXL embeddings.

Licencja: NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.

Testy bezpieczeństwa: NVIDIA applied dataset-level filters for unsafe content and faces, plus output filters and watermarking guidance. The model is positioned as a research artifact and not a deployed user-facing product, so no formal RSP-style policy is published.

Znane ograniczenia

  • Short rollouts (a few seconds) before drift dominates
  • Not a controller - cannot produce robot actions on its own
  • Heavy GPU footprint for the 14B variant
  • Domain skew toward driving / Western indoor scenes
  • Output is video, not joint commands - needs downstream policy
  • Restricted commercial license (Open Model License, not Apache)
03

Ceny

Obecnie niedostępne. Dla tego modelu nie ma ceny w tej chwili, dlatego nie można go uruchomić.

04

API

Wywołaj NVIDIA Cosmos-Predict-1 za pomocą klucza API Railwail. Użyj tego ID modelu w żądaniu:

Niedostępne przez API

Modele robotyki działają na sprzęcie robotycznym, a nie przez API railwail.

05

Specyfikacje

ID modelu
cosmos-predict-1
Deweloper
Other
Wejście
Tekst, Obraz
Wyjście
Akcje robota
Rozmiar modelu
7B and 14B variants (Predict-1)
Licencja
NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.
Wpis w katalogu zaktualizowany
24 września 2026

Tagi

  • nvidia
  • cosmos
  • vla
  • robotics
  • research-only
  • open-weights
  • world-model
06

Przypadki użycia

Do czego się go używa

  • Synthetic data generation for robotics and AV
  • World-model research and benchmarks
  • Offline policy evaluation in imagined rollouts
  • Simulation-to-real video adaptation pipelines
  • Pretraining backbone for VLA models
  • Academic embodied-AI research (research-only license)
07

Często zadawane pytania

Co to jest NVIDIA Cosmos-Predict-1?

NVIDIA Cosmos-Predict-1 to model opracowany przez Other w kategorii Robotyka / VLA. Jest wymieniony w katalogu Railwail, ale nie może być uruchomiony w tej chwili.

Ile kosztuje NVIDIA Cosmos-Predict-1 w serwisie Railwail?

NVIDIA Cosmos-Predict-1 nie może być uruchomiony w serwisie Railwail w tej chwili, dlatego nie ma aktualnej ceny. Dostępne alternatywy z cenami są wymienione poniżej na tej stronie.

Jak szybki jest NVIDIA Cosmos-Predict-1?

Dla NVIDIA Cosmos-Predict-1 jest jeszcze zbyt mało zmierzonych przebiegów w serwisie Railwail, aby podać czas przebiegu. Zależy to od wejścia, ustawień i obciążenia u dostawcy.

Kiedy powinienem użyć NVIDIA Cosmos-Predict-1?

NVIDIA Cosmos-Predict-1 należy do kategorii Robotyka / VLA. Strona kategorii zawiera listę innych modeli tego typu z ich cenami.

Wszystkie modele: Robotyka / VLA

Czy NVIDIA Cosmos-Predict-1 może przetwarzać obrazy?

Tak. NVIDIA Cosmos-Predict-1 akceptuje obrazy jako wejście oprócz tekstu.

Czy mogę używać NVIDIA Cosmos-Predict-1 teraz?

Obecnie niedostępne. Strona pozostaje online; dostępne alternatywy z tej samej kategorii są wymienione poniżej.

Wszystkie modele przez jedno API

Jeden klucz API dla każdego modelu na Railwail. Opłaty pobierane są z przedpłaconych kredytów, 1 kredyt = 0,01 USD.