Grounded-SAM
grounded-samGrounding DINO plus SAM. Open-vocabulary text-prompted detection and segmentation in one pipeline for fully-automatic mask generation.
- Price
- โ $0.0015/run
- Input โ output
- Text + Image โ Text
- Developer
- Community
- Updated
- September 23, 2026
Playground
Try Grounded-SAM
No input form
No input form for this model yet
Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.
Examples
Open full sizePrompt
clothes,shoes
Open full size
InputPrompt
Car
About Grounded-SAM
Grounded-SAM is a model by Community in the Multimodal category. On Railwail, Grounded-SAM costs โ $0.0015 per run.
Pricing
| Typical run (โ 1 s on L40S) | $0.0015 per run |
|---|---|
| GPU time (L40S) | $0.00117 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3ร the typical price is reserved from your balance and settled afterwards.
- 1 credit = $0.01
Cost calculator
Price calculator
Typical according to the provider: about 1.2 s
Total
$0.15
15 credits
Per run
$0.0015 ยท 0.15 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
grounded-sam- Developer
- Community
- Category
- Multimodal
- Input
- Text, Image
- Output
- Text
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- September 23, 2026
Tags
- replicate
- segmentation
- vision-understanding
- open-vocabulary
- open-source
Frequently asked questions
What is Grounded-SAM?
Grounded-SAM is a model by Community in the Multimodal category.
How much does Grounded-SAM cost on Railwail?
On Railwail, Grounded-SAM costs โ $0.0015 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.
How fast is Grounded-SAM?
There are not enough measured runs of Grounded-SAM on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Grounded-SAM better than BLIP?
That depends on the task. Grounded-SAM (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare Grounded-SAM and BLIPCan Grounded-SAM process images?
Yes. Grounded-SAM accepts images as input in addition to text.
Comparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.