NVIDIA Cosmos-Predict-1

Robotik / VLAMevcut Değil
Other tarafındanModel Kimliği: cosmos-predict-1

NVIDIA's world foundation model for physical AI. Diffusion-based video prediction for robotics simulation.

Durum
Mevcut Değil
Giriş → Çıkış
Metin + Görüntü → Robot İşlemleri
Geliştirici
Other
Güncellendi
24 Eylül 2026

NVIDIA Cosmos-Predict-1 şu anda kullanılamıyor

Bu sayfadaki ayrıntıları yine de okuyabilirsiniz. Hemen karşılaştırılabilir bir modeli çalıştırmak için aşağıdaki mevcut alternatiflerden birini seçin.

01

Playground

NVIDIA Cosmos-Predict-1

Araştırma modeli

Şu anda kullanılamıyor

NVIDIA Cosmos-Predict-1 bir robotik modeldir (vision-language-action) ve railwail API'si aracılığıyla çalıştırılamaz.

02

NVIDIA Cosmos-Predict-1 Hakkında

Özet24 Eylül 2026 itibariyle

NVIDIA Cosmos-Predict-1, Other tarafından Robotik / VLA kategorisinde geliştirilen bir modeldir. NVIDIA Cosmos-Predict-1 şu anda Railwail üzerinde kullanılamıyor.

Arka plan

NVIDIA hakkında

Kuruluş yılı 1993 · Santa Clara, California, USA

NVIDIA is the dominant supplier of GPUs for AI training and inference and runs a large in-house research organisation across robotics, simulation, and generative modelling. NVIDIA Cosmos was announced at CES 2025 as a family of 'World Foundation Models' (WFMs) for Physical AI - models that predict how the physical world evolves given video, language, and action conditioning. Cosmos is positioned as a developer platform for robotics and autonomous-vehicle teams to generate synthetic training data, run policy evaluations in simulation, and bootstrap Vision-Language-Action (VLA) pipelines. The 'Predict-1' track focuses on diffusion-based video-future prediction conditioned on text and/or first-frame inputs and ships in 7B and 14B parameter variants with open weights under the NVIDIA Open Model License.

NVIDIA ziyaret edin

Mimari

Diffusion-based world foundation model (text/video-to-video) for Physical AI

Cosmos-Predict-1 is a diffusion world model that predicts future video frames conditioned on text prompts, a starting frame, or short context clips. It uses a 3D causal video tokenizer (Cosmos Tokenizer) to compress video into spatio-temporal latents, then runs a Diffusion Transformer in latent space with cross-attention to text embeddings produced by a T5-XXL encoder. Training data is a curated corpus of ~20 million hours of driving, robotics, and human-activity video, filtered for motion quality, captioning coverage and safety. The model is not itself a VLA controller, but is the world-model backbone of NVIDIA's Cosmos stack: Cosmos-Predict generates rollouts; Cosmos-Reason adds VLM reasoning over predicted futures; and Cosmos-Transfer adapts simulation-to-real video. In a VLA pipeline it provides synthetic 'imagined' trajectories and dense reward / value signals, and is used to evaluate manipulation and driving policies offline at scale.

Parametreler
7B and 14B variants (Predict-1)

Yetenekler

  • Predicts future video conditioned on text, image, or video context
  • Two open-weight variants: Predict-1-7B and Predict-1-14B
  • Generates physically plausible motion for driving, manipulation, and humanoid scenes
  • Integrates with NVIDIA Isaac, Omniverse, and DRIVE pipelines
  • Used as synthetic data engine for VLA / autonomy training
  • Supports prompt upsampling via Cosmos-Reason VLM
  • Cosmos Tokenizer (3D causal VAE) can be reused as a video encoder
  • Released alongside Cosmos-Reason and Cosmos-Transfer for full Physical AI stack
  • Best for: synthetic data, world-model research, robotics simulation, AV training.

Eğitim ve lisans

Trained on ~20 million hours of curated physical-world video (driving, robotics manipulation, humanoid / first-person, navigation) sourced from licensed and open datasets, with multi-stage filtering for motion quality, caption alignment and safety. Text conditioning uses T5-XXL embeddings.

Lisans: NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.

Güvenlik testleri: NVIDIA applied dataset-level filters for unsafe content and faces, plus output filters and watermarking guidance. The model is positioned as a research artifact and not a deployed user-facing product, so no formal RSP-style policy is published.

Bilinen sınırlamalar

  • Short rollouts (a few seconds) before drift dominates
  • Not a controller - cannot produce robot actions on its own
  • Heavy GPU footprint for the 14B variant
  • Domain skew toward driving / Western indoor scenes
  • Output is video, not joint commands - needs downstream policy
  • Restricted commercial license (Open Model License, not Apache)
03

Fiyatlandırma

Şu anda kullanılamıyor. Bu modelin şu anda bir fiyatı yok, bu nedenle çalıştırılamıyor.

04

API

NVIDIA Cosmos-Predict-1 öğesini Railwail API anahtarınızla çağırın. İstekte bu model kimliğini kullanın:

API aracılığıyla kullanılamaz

Robotik modeller railwail API aracılığıyla değil, robot donanımında çalışır.

05

Özellikler

Model Kimliği
cosmos-predict-1
Geliştirici
Other
Giriş
Metin, Görüntü
Çıkış
Robot İşlemleri
Model boyutu
7B and 14B variants (Predict-1)
Lisans
NVIDIA Open Model License - research-only / developer use with restrictions; weights downloadable from Hugging Face and NGC.
Katalog girişi güncellendi
24 Eylül 2026

Etiketler

  • nvidia
  • cosmos
  • vla
  • robotics
  • research-only
  • open-weights
  • world-model
06

Kullanım Alanları

Nerelerde Kullanılır

  • Synthetic data generation for robotics and AV
  • World-model research and benchmarks
  • Offline policy evaluation in imagined rollouts
  • Simulation-to-real video adaptation pipelines
  • Pretraining backbone for VLA models
  • Academic embodied-AI research (research-only license)
07

Sık Sorulan Sorular

NVIDIA Cosmos-Predict-1 nedir?

NVIDIA Cosmos-Predict-1, Other tarafından Robotik / VLA kategorisinde oluşturulan bir modeldir. Railwail'de listelenmiştir ancak şu anda çalıştırılamaz.

NVIDIA Cosmos-Predict-1 Railwail'de ne kadar maliyetlidir?

NVIDIA Cosmos-Predict-1 şu anda Railwail'de çalıştırılamaz, bu nedenle mevcut bir fiyat yoktur. Fiyatları olan mevcut alternatifler bu sayfanın aşağısında listelenmiştir.

NVIDIA Cosmos-Predict-1 ne kadar hızlıdır?

Railwail'de NVIDIA Cosmos-Predict-1 için henüz çalıştırma süresi belirtmek için yeterli ölçülen çalıştırma yoktur. Bu, giriş, ayarlar ve sağlayıcıdaki yüke bağlıdır.

NVIDIA Cosmos-Predict-1'yi ne zaman kullanmalıyım?

NVIDIA Cosmos-Predict-1, Robotik / VLA kategorisine aittir. Kategori sayfası bu türün diğer modellerini fiyatlarıyla birlikte listeler.

Tüm modeller: Robotik / VLA

NVIDIA Cosmos-Predict-1 görüntüleri işleyebilir mi?

Evet. NVIDIA Cosmos-Predict-1 metne ek olarak giriş olarak görüntüleri kabul eder.

NVIDIA Cosmos-Predict-1'yi şu anda kullanabilir miyim?

Şu anda kullanılamıyor. Sayfa çevrimiçi kalır; aynı kategoriden mevcut alternatifler aşağıda listelenmiştir.

Tüm Modeller Tek Bir API Aracılığıyla

Railwail'deki her model için bir API anahtarı. Kullanım, ön ödemeli kredilerden tahsil edilir, 1 kredi = $0,01.