$1.5121 (151.21 credits) are reserved at the start; the actual GPU time is billed.
For accounts without a purchase: runs above 2 credits need a top-up.
New here?
10 free credits ($0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.
02
Examples
Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
Prompt
Close-up of a chameleon's eye, with its scaly skin changing color. Ultra high resolution 4k.
Length: 0:05
Settings
num_frames
121
fps
24
Prompt
A pristine snowglobe featuring a winter scene sits peacefully. The globe violently explodes, sending glass, water, and glittering fake snow in all directions. The scene is captured with high-speed photography.
Length: 0:05
Settings
num_frames
121
fps
24
Prompt
The video opens with a close-up of a woman in a white and purple outfit, holding a glowing purple butterfly. She has dark hair and walks gracefully through a traditional Japanese-style village at night
Length: 0:05
Settings
num_frames
121
fps
24
03
About Mochi 1
TL;DRAs of September 23, 2026
Mochi 1 is a model by Community in the Video generation category. On Railwail, Mochi 1 costs β $0.5041 per run. Supported aspect ratios: 16:9, 9:16, and 1:1.
Background
About Genmo
Founded 2023 Β· San Francisco, USA
Genmo was founded in 2023 by Paras Jain (CEO) and Ajay Jain (CTO), both PhDs from UC Berkeley's BAIR lab, with a focus on open-source generative video. The company released the early Replay product and the smaller GEN-1 video model before launching Mochi 1 in October 2024 as a 10B-parameter open-weight text-to-video model under the Apache 2.0 licence -- at the time the largest open-source video model with a fully permissive licence. Genmo positions Mochi 1 as a research foundation for the community to fine-tune, extend and study, in deliberate contrast to closed competitors. The company has raised over $30M from investors including NEA and the Y Combinator network.
Asymmetric Diffusion Transformer (AsymmDiT) at 10B parameters with custom 3D VAE
Mochi 1 is a 10B-parameter Asymmetric Diffusion Transformer (AsymmDiT) operating on a high-compression 3D causal Variational Autoencoder. The 'asymmetric' design uses dramatically more parameters for the video stream than for the text stream while sharing self-attention, on the hypothesis that visual modeling is the bottleneck for video generation. Position information uses 3D RoPE; the model is trained with Rectified Flow Matching at full resolution and high motion intensity. Text conditioning uses a T5-XXL encoder. Mochi 1 generates 5-second clips at 480p / 30 fps natively (with a 720p HD variant in preview at launch). Training is done on a curated, filtered video corpus with dense captions produced by an in-house captioner. The team explicitly report ablations on resolution scheduling, motion intensity filtering and caption quality.
Parameters
10 billion
Capabilities
10B open-weight text-to-video model under Apache 2.0 (most permissive in class)
Asymmetric DiT design biased toward visual capacity
curl https://railwail.com/api/v1/videos/generations \
-H "Authorization: Bearer $RAILWAIL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "mochi-1-genmo",
"prompt": "A slow drone shot over a misty pine forest at sunrise"
}'
# Response: {"job_id": "...", "status": "queued", ...}
# Poll until status is completed, failed or cancelled:
curl https://railwail.com/api/v1/jobs/JOB_ID \
-H "Authorization: Bearer $RAILWAIL_API_KEY"
import os
import time
import requests
API = "https://railwail.com/api/v1"
headers = {"Authorization": f"Bearer {os.environ['RAILWAIL_API_KEY']}"}
job = requests.post(
f"{API}/videos/generations",
headers=headers,
json={
"model": "mochi-1-genmo",
"prompt": "A slow drone shot over a misty pine forest at sunrise",
},
).json()
while True:
status = requests.get(f"{API}/jobs/{job['job_id']}", headers=headers).json()
if status["status"] in ("completed", "failed", "cancelled"):
break
time.sleep(5)
print(status["status"], status.get("output_url"))
const API = "https://railwail.com/api/v1";
const headers = {
Authorization: `Bearer ${process.env.RAILWAIL_API_KEY}`,
"Content-Type": "application/json",
};
const job = await fetch(`${API}/videos/generations`, {
method: "POST",
headers,
body: JSON.stringify({
model: "mochi-1-genmo",
prompt: "A slow drone shot over a misty pine forest at sunrise"
}),
}).then((r) => r.json());
let status;
do {
await new Promise((r) => setTimeout(r, 5000));
status = await fetch(`${API}/jobs/${job.job_id}`, { headers }).then((r) => r.json());
} while (!["completed", "failed", "cancelled"].includes(status.status));
console.log(status.status, status.output_url);
Mochi 1 is a model by Community in the Video generation category. On Railwail you can call it with an API key through the Railwail API.
How much does Mochi 1 cost on Railwail?
On Railwail, Mochi 1 costs β $0.5041 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.
Which settings does Mochi 1 support?
According to its input schema, Mochi 1 knows these parameters: prompt (up to 2,000 characters), fps (8 to 30), seed, aspect_ratio (16:9, 9:16, or 1:1), duration_sec (1 to 6), and motion_strength (0 to 1).
How fast is Mochi 1?
There are not enough measured runs of Mochi 1 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Mochi 1 better than Google Veo 3.1?
That depends on the task. Mochi 1 (Community) and Google Veo 3.1 (Google DeepMind) are both models in the Video generation category. The comparison page shows their prices and specifications side by side.
Create a Railwail API key and send your request with the model ID mochi-1-genmo. Code examples for curl, Python and JavaScript are in the API section of this page.
Tencent's HunyuanVideo, a 13B open-weights text-to-video diffusion transformer. Produces high-motion, photorealistic clips with smooth temporal consistency and was one of the first open models to rival closed systems on motion quality.