Molmo 7B
molmo-7bAllen AI Molmo 7B-D on Replicate. Open vision-language model trained on the PixMo data, notable for pointing at and locating objects in images, not just describing them.
- Price
- β $0.0505/run
- Input β output
- Text + Image β Text
- Developer
- Community
- Updated
- September 23, 2026
Playground
Try Molmo 7B
Input & output
This run
about $0.0505 Β· 5.05 credits
$0.1513 (15.13 credits) are reserved at the start; the actual GPU time is billed.
For accounts without a purchase: runs above 2 credits need a top-up.
New here?
10 free credits ($0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.
Examples
InputPrompt
What do you see? Give me a detailed answer
Response
I see a charming scene featuring a large, fluffy white dog sitting on a wooden bench in the middle of a field. The dog appears to be a poodle mix, with curly fur covering its entire body. It's sitting upright with its front paws hanging over the edge of the bench, looking directly at the camera with a happy expression. The dog's tongue is hanging out, and its eyes are dark and expressive. The bench is made of dark brown wood with a lattice design on the backrest. It's positioned on a patch of brown, dead grass, which suggests it might be late fall or early winter. In the background, there's a vast field filled with tall, brown grass, and further back, I can see a line of bare trees against a light blue sky. The overall atmosphere is peaceful and serene. The dog seems to be enjoying its time outdoors, perhaps waiting for its owner or just taking in the scenery. The contrast between the white, fluffy dog and the brown
About Molmo 7B
Molmo 7B is a model by Community in the Multimodal category. On Railwail, Molmo 7B costs β $0.0505 per run.
Pricing
| Typical run (β 43 s on L40S) | $0.0505 per run |
|---|---|
| GPU time (L40S) | $0.00117 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3Γ the typical price is reserved from your balance and settled afterwards.
- 1 credit = $0.01
Cost calculator
Price calculator
Typical according to the provider: about 43.1 s
Total
$5.05
505 credits
Per run
$0.0505 Β· 5.05 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
molmo-7b- Developer
- Community
- Category
- Multimodal
- Input
- Text, Image
- Output
- Text
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- September 23, 2026
Input parameters
Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.
promptrequiredQuestion about the image
Type: TextDefault: βAllowed values: up to 16,000 charactersimage_urlImage URL to analyze
Type: TextDefault: βAllowed values: βmax_tokensType: IntegerDefault:1024Allowed values: 1 to 4,096
Tags
- replicate
- allenai
- molmo
- vision-understanding
- open-weights
- grounding
Frequently asked questions
What is Molmo 7B?
Molmo 7B is a model by Community in the Multimodal category.
How much does Molmo 7B cost on Railwail?
On Railwail, Molmo 7B costs β $0.0505 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.
Which settings does Molmo 7B support?
According to its input schema, Molmo 7B knows these parameters: prompt (up to 16,000 characters), image_url, and max_tokens (1 to 4,096).
How fast is Molmo 7B?
There are not enough measured runs of Molmo 7B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Molmo 7B better than BLIP?
That depends on the task. Molmo 7B (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare Molmo 7B and BLIPCan Molmo 7B process images?
Yes. Molmo 7B accepts images as input in addition to text.
Comparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.