Moondream2
moondream2Moondream2 small vision-language model on Replicate. About 1.9B params, designed to run on edge devices, handles captioning, visual QA and short OCR-style reads at very low cost.
- Price
- โ $0.0020/run
- Input โ output
- Text + Image โ Text
- Developer
- Community
- Updated
- September 23, 2026
Playground
Try Moondream2
Input & output
This run
about $0.002 ยท 0.2 credits
$0.0058 (0.58 credits) are reserved at the start; the actual GPU time is billed.
New here?
10 free credits ($0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits. Enough for 17 runs of this model.
Examples
InputPrompt
Describe this image
Response
The image features a logo with a smiling blue circle above the word "moondream" written in black text.
InputPrompt
Describe this image
Response
A man with a beard and mustache, wearing a suit and red tie, is smiling at the camera with a blue background featuring a logo.
InputPrompt
Describe this image
Response
An astronaut in a white spacesuit is riding a unicorn with a rainbow mane and tail, soaring through a colorful, dreamlike sky with clouds and rainbows.
About Moondream2
Moondream2 is a model by Community in the Multimodal category. On Railwail, Moondream2 costs โ $0.0020 per run.
Pricing
| Typical run (โ 2 s on L40S) | $0.0020 per run |
|---|---|
| GPU time (L40S) | $0.00117 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3ร the typical price is reserved from your balance and settled afterwards.
- 1 credit = $0.01
Cost calculator
Price calculator
Typical according to the provider: about 1.6 s
Total
$0.20
20 credits
Per run
$0.002 ยท 0.2 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
moondream2- Developer
- Community
- Category
- Multimodal
- Input
- Text, Image
- Output
- Text
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- September 23, 2026
Input parameters
Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.
promptrequiredQuestion about the image
Type: TextDefault: โAllowed values: up to 8,000 charactersimage_urlImage URL to analyze
Type: TextDefault: โAllowed values: โ
Tags
- replicate
- moondream
- vision-understanding
- open-source
- small
- edge
Frequently asked questions
What is Moondream2?
Moondream2 is a model by Community in the Multimodal category.
How much does Moondream2 cost on Railwail?
On Railwail, Moondream2 costs โ $0.0020 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.
Which settings does Moondream2 support?
According to its input schema, Moondream2 knows these parameters: prompt (up to 8,000 characters) and image_url.
How fast is Moondream2?
There are not enough measured runs of Moondream2 on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Moondream2 better than BLIP?
That depends on the task. Moondream2 (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare Moondream2 and BLIPCan Moondream2 process images?
Yes. Moondream2 accepts images as input in addition to text.
Comparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.