CLIP Interrogator
clip-interrogatorpharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Price
- β US$0.0457/run
- Input β output
- Text + Image β Text
- Developer
- Community
- Updated
- 23 September 2026
Playground
Try CLIP Interrogator
Input & output
This run
about US$0.0457 Β· 4.57 credits
US$0.1369 (13.69 credits) are reserved at the start; the actual GPU time is billed.
For accounts without a purchase: runs above 2 credits need a top-up.
New here?
10 free credits (US$0.10) when you sign up with Google
Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits.
Examples
InputNo text prompt: the model only takes the input shown.
Response
a watercolor painting of a sea turtle, a digital painting, by Kubisi art, featured on dribbble, medibang, warm saturated palette, red and green tones, turquoise horizon, digital art h 9 6 0, detailed scenery βwidth 672, illustration:.4, spray art, artstatiom
InputNo text prompt: the model only takes the input shown.
Response
a painting of a field with green grass, by Vincent Van Gogh, featured on deviantart, windy day, virtuosic level detail, the front of a trading card, loosely cropped, expansive, tendrils in the background, photo courtesy museum of art, 1980s art, standing on a hill, creative commons attribution
About CLIP Interrogator
CLIP Interrogator is a model by Community in the Multimodal category. On Railwail, CLIP Interrogator costs β US$0.0457 per run.
Pricing
| Typical run (β 169 s on T4) | US$0.0457 per run |
|---|---|
| GPU time (T4) | US$0.00027 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3Γ the typical price is reserved from your balance and settled afterwards.
- 1 credit = US$0.01
Cost calculator
Price calculator
Typical according to the provider: about 168.9 s
Total
US$4.57
457 credits
Per run
US$0.0457 Β· 4.57 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
clip-interrogator- Developer
- Community
- Category
- Multimodal
- Input
- Text, Image
- Output
- Text
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- 23 September 2026
Input parameters
Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.
imagerequiredImage to interrogate
Type: TextDefault: βAllowed values: βmodeType: ChoiceDefault:bestAllowed values: best, fast, classic or negativeclip_model_nameType: TextDefault:ViT-L-14/openaiAllowed values: β
Tags
- replicate
- clip-interrogator
- captioning
- tagging
- clip
- blip
- prompt-generation
- image
Frequently asked questions
What is CLIP Interrogator?
CLIP Interrogator is a model by Community in the Multimodal category.
How much does CLIP Interrogator cost on Railwail?
On Railwail, CLIP Interrogator costs β US$0.0457 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.
Which settings does CLIP Interrogator support?
According to its input schema, CLIP Interrogator knows these parameters: image, mode (best, fast, classic or negative) and clip_model_name.
How fast is CLIP Interrogator?
There are not enough measured runs of CLIP Interrogator on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is CLIP Interrogator better than BLIP?
That depends on the task. CLIP Interrogator (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare CLIP Interrogator and BLIPCan CLIP Interrogator process images?
Yes. CLIP Interrogator accepts images as input in addition to text.
Comparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
Meta Segment Anything 2. Promptable segmentation across images and video with temporal memory. Zero-shot, point/box/mask prompts, fast on a single H100.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.