La start se rezervă 1,5121 USD (151,21 credite); se factureaza timpul GPU real.
Pentru conturi fără achiziție anterioară: rulările peste 2 credite necesită o reîncărcare.
Nou aici?
10 credite gratuite (0,10 USD) când te înregistrezi cu Google
Utilizabil 24 ore după înregistrare, până la 5 rulări pe zi și maximum 2 credite pe rulare. Alte metode de conectare încep fără credite.
02
Examples
Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
Prompt
Close-up of a chameleon's eye, with its scaly skin changing color. Ultra high resolution 4k.
Lungime: 0:05
Settings
num_frames
121
fps
24
Prompt
A pristine snowglobe featuring a winter scene sits peacefully. The globe violently explodes, sending glass, water, and glittering fake snow in all directions. The scene is captured with high-speed photography.
Lungime: 0:05
Settings
num_frames
121
fps
24
Prompt
The video opens with a close-up of a woman in a white and purple outfit, holding a glowing purple butterfly. She has dark hair and walks gracefully through a traditional Japanese-style village at night
Lungime: 0:05
Settings
num_frames
121
fps
24
03
Despre Mochi 1
Pe scurtDin 23 septembrie 2026
Mochi 1 este un model de Community din categoria Generare videoclipuri. Pe Railwail, Mochi 1 costă ≈ 0,5041 USD per rulare. Rapoarte de aspect acceptate: 16:9, 9:16 și 1:1.
Fundal
Despre Genmo
Fondat 2023 · San Francisco, USA
Genmo was founded in 2023 by Paras Jain (CEO) and Ajay Jain (CTO), both PhDs from UC Berkeley's BAIR lab, with a focus on open-source generative video. The company released the early Replay product and the smaller GEN-1 video model before launching Mochi 1 in October 2024 as a 10B-parameter open-weight text-to-video model under the Apache 2.0 licence -- at the time the largest open-source video model with a fully permissive licence. Genmo positions Mochi 1 as a research foundation for the community to fine-tune, extend and study, in deliberate contrast to closed competitors. The company has raised over $30M from investors including NEA and the Y Combinator network.
Asymmetric Diffusion Transformer (AsymmDiT) at 10B parameters with custom 3D VAE
Mochi 1 is a 10B-parameter Asymmetric Diffusion Transformer (AsymmDiT) operating on a high-compression 3D causal Variational Autoencoder. The 'asymmetric' design uses dramatically more parameters for the video stream than for the text stream while sharing self-attention, on the hypothesis that visual modeling is the bottleneck for video generation. Position information uses 3D RoPE; the model is trained with Rectified Flow Matching at full resolution and high motion intensity. Text conditioning uses a T5-XXL encoder. Mochi 1 generates 5-second clips at 480p / 30 fps natively (with a 720p HD variant in preview at launch). Training is done on a curated, filtered video corpus with dense captions produced by an in-house captioner. The team explicitly report ablations on resolution scheduling, motion intensity filtering and caption quality.
Parametri
10 billion
Capabilități
10B open-weight text-to-video model under Apache 2.0 (most permissive in class)
Asymmetric DiT design biased toward visual capacity
curl https://railwail.com/api/v1/videos/generations \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mochi-1-genmo",
"prompt": "A slow drone shot over a misty pine forest at sunrise"
}'
# Response: {"job_id": "...", "status": "queued", ...}
# Poll until status is completed, failed or cancelled:
curl https://railwail.com/api/v1/jobs/JOB_ID \
-H "Authorization: Bearer $RAILWAIL_API_KEY"
import os
import time
import requests
API = "https://railwail.com/api/v1"
headers = {"Authorization": f"Bearer {os.environ['RAILWAIL_API_KEY']}"}
job = requests.post(
f"{API}/videos/generations",
headers=headers,
json={
"model": "mochi-1-genmo",
"prompt": "A slow drone shot over a misty pine forest at sunrise",
},
).json()
while True:
status = requests.get(f"{API}/jobs/{job['job_id']}", headers=headers).json()
if status["status"] in ("completed", "failed", "cancelled"):
break
time.sleep(5)
print(status["status"], status.get("output_url"))
const API = "https://railwail.com/api/v1";
const headers = {
Authorization: `Bearer ${process.env.RAILWAIL_API_KEY}`,
"Content-Type": "application/json",
};
const job = await fetch(`${API}/videos/generations`, {
method: "POST",
headers,
body: JSON.stringify({
model: "mochi-1-genmo",
prompt: "A slow drone shot over a misty pine forest at sunrise"
}),
}).then((r) => r.json());
let status;
do {
await new Promise((r) => setTimeout(r, 5000));
status = await fetch(`${API}/jobs/${job.job_id}`, { headers }).then((r) => r.json());
} while (!["completed", "failed", "cancelled"].includes(status.status));
console.log(status.status, status.output_url);
Mochi 1 este un model de Community din categoria Generare videoclipuri. Pe Railwail îl poți apela cu o cheie API prin Railwail API.
Cât costă Mochi 1 pe Railwail?
Pe Railwail, Mochi 1 costă ≈ 0,5041 USD per rulare. Ți se percepe taxa pentru ceea ce fiecare cerere folosește efectiv. Utilizarea se plătește din credite prepay; 1 credit egal cu 0,01 USD.
Ce setări acceptă Mochi 1?
Conform schemei sale de intrare, Mochi 1 cunoaște acești parametri: prompt (până la 2.000 caractere), fps (8 până la 30), seed, aspect_ratio (16:9, 9:16 sau 1:1), duration_sec (1 până la 6) și motion_strength (0 până la 1).
Cât de rapid este Mochi 1?
Nu sunt suficiente rulări măsurate ale Mochi 1 pe Railwail încă pentru a indica un timp de rulare. Depinde de intrare, de setări și de sarcina la furnizor.
Este Mochi 1 mai bun decât Google Veo 3.1?
Depinde de sarcină. Mochi 1 (Community) și Google Veo 3.1 (Google DeepMind) sunt ambele modele din categoria Generare videoclipuri. Pagina de comparație arată prețurile și specificațiile lor una lângă alta.
Creează o cheie Railwail API și trimite cererea cu ID-ul modelului mochi-1-genmo. Exemple de cod pentru curl, Python și JavaScript sunt în secțiunea API a acestei pagini.
Tencent's HunyuanVideo, a 13B open-weights text-to-video diffusion transformer. Produces high-motion, photorealistic clips with smooth temporal consistency and was one of the first open models to rival closed systems on motion quality.