Moondream2
moondream2Moondream2 small vision-language model on Replicate. About 1.9B params, designed to run on edge devices, handles captioning, visual QA and short OCR-style reads at very low cost.
- Pris
- ≈ 0,0020 US$/kørsel
- Input → output
- Tekst + Billede → Tekst
- Udvikler
- Community
- Opdateret
- 23. september 2026
Playground
Prøv Moondream2
Input & resultat
Denne kørsel
ca. 0,002 US$ · 0,2 credits
0,0058Â US$ (0,58 credits) reserveres ved start; den faktiske GPU-tid faktureres.
Ny her?
10 gratis credits (0,10 US$) når du tilmelder dig med Google
Kan bruges 24 timer efter tilmelding, op til 5 kørsler pr. dag og højst 2 credits pr. kørsel. Andre login-metoder starter uden credits. Nok til 17 kørsler af denne model.
Examples
InputPrompt
Describe this image
Response
The image features a logo with a smiling blue circle above the word "moondream" written in black text.
InputPrompt
Describe this image
Response
A man with a beard and mustache, wearing a suit and red tie, is smiling at the camera with a blue background featuring a logo.
InputPrompt
Describe this image
Response
An astronaut in a white spacesuit is riding a unicorn with a rainbow mane and tail, soaring through a colorful, dreamlike sky with clouds and rainbows.
Om Moondream2
Moondream2 er en model af Community i kategorien Multimodal. På Railwail koster Moondream2 ≈ 0,0020 US$ pr. kørsel.
Priser
| Typisk kørsel (≈ 2 s på L40S) | 0,0020 US$ pr. kørsel |
|---|---|
| GPU-tid (L40S) | 0,00117Â US$ pr. GPU-sekund |
- Faktureres efter den GPU-tid, som kørslen faktisk tager. Når kørslen starter, reserveres 3× den typiske pris fra din saldo og afregnes efterfølgende.
- 1 kredit = 0,01Â US$
Omkostningsberegner
Prisberegner
Typisk ifølge udbyderen: ca. 1,6 s
I alt
0,20Â US$
20 credits
Pr. kørsel
0,002 US$ · 0,2 credits
Fakturering efter faktisk GPU-tid; dette er et estimat.
API
Intet bekræftet API-eksempel
Den offentlige API sender et andet inputformat end det, som denne model har brug for. Brug playground ovenfor.
Specifikationer
- Model-ID
moondream2- Udvikler
- Community
- Kategori
- Multimodal
- Input
- Tekst, Billede
- Output
- Tekst
- Fakturering
- Efter forbrug (tokens eller GPU-tid)
- Katalogelement opdateret
- 23. september 2026
Inputparametre
Inputs og indstillinger fra modellens inputskema. Eksemplet i API-afsnittet viser, hvilke af dem API'en accepterer.
promptpåkrævetQuestion about the image
Type: TekstStandard: –Tilladte værdier: op til 8.000 tegnimage_urlImage URL to analyze
Type: TekstStandard: –Tilladte værdier: –
Tags
- replicate
- moondream
- vision-understanding
- open-source
- small
- edge
Ofte stillede spørgsmål
Hvad er Moondream2?
Moondream2 er en model fra Community i kategorien Multimodal.
Hvad koster Moondream2 på Railwail?
På Railwail koster Moondream2 ≈ 0,0020 US$ pr. kørsel. Du betaler for det, som hver anmodning faktisk bruger. Forbrug betales fra forudbetalte credits; 1 credit svarer til 0,01 US$.
Hvilke indstillinger understøtter Moondream2?
Ifølge dens inputskema kender Moondream2 disse parametre: prompt (op til 8.000 tegn) og image_url.
Hvor hurtig er Moondream2?
Der er endnu ikke nok målte kørsler af Moondream2 på Railwail til at angive en udførelsestid. Det afhænger af inputtet, indstillingerne og belastningen hos provideren.
Er Moondream2 bedre end BLIP?
Det afhænger af opgaven. Moondream2 (Community) og BLIP (Salesforce) er begge modeller i kategorien Multimodal. Sammenligningssiden viser deres priser og specifikationer side om side.
Sammenlign Moondream2 og BLIPKan Moondream2 behandle billeder?
Ja. Moondream2 accepterer billeder som input ud over tekst.
Sammenlignelige modeller
Alle i denne kategori- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
Alle modeller via én API
En API-nøgle til alle modeller på Railwail. Forbrug debiteres fra forudbetalte credits, 1 credit = 0,01 US$.