Idefics3 8B
idefics3-8bHugging Face Idefics3 8B. Llama-3 based open-source vision-language model with strong document QA and chart-understanding performance.
- Price
- โ $0.0012/run
- Context
- 8,192 tokens
- Max. output
- 4,096 tokens
- Input โ output
- Text + Image โ Text
- Developer
- Community
- Updated
- September 23, 2026
Playground
Try Idefics3 8B
No input form
No input form for this model yet
Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.
Examples
InputPrompt
What do you see? Give me a detailed answer
Response
A white dog is sitting on the bench. The background of the image is blurred, but we can still see trees and dry grass in the background. There are clouds visible in the sky.
InputPrompt
What do you see? Give me a detailed answer
Response
The image depicts an old-fashioned typewriter that has been modified to resemble the face of a robot. This is achieved by adding various mechanical components and parts, transforming it into what appears as a humanoid visage with distinct features. The top part of the typewriter's body forms the head section, which includes two large circular eyes made from gears or other metallic objects. These eyes are positioned symmetrically on either side of the centerline, giving the impression of staring directly at the viewer. Each eye contains intricate details such as small metal pieces arranged in a way that mimics human-like expressions, possibly suggesting surprise or curiosity. Below these "eyes," there is a prominent nose formed by a series of horizontal bars connected vertically, creating a bridge-like structure between them. Below this nasal area lies the mouth, which consists of several rows of keys from the original typewriter keyboard. These keys have been painted white for contrast against the dark background, making them stand out prominently. On each row, the letters QWERTY (typical of standard English keyboards) can be seen clearly visible. However, some additional characters like U, P, Y, I, W, E, R, T, S, D, F, G, H, J, K, L, Z, X, C, V, B, N, M appear to be present, indicating that they might belong to different languages or variations of typing layouts. To complete the robotic appearance, the entire assembly is framed within a rectangular border resembling aโฆ
About Idefics3 8B
Idefics3 8B is a model by Community in the Multimodal category. On Railwail, Idefics3 8B costs โ $0.0012 per run. The context window holds 8,192 tokens, and one response can be up to 4,096 tokens long.
Pricing
| Typical run (โ 1 s on L40S) | $0.0012 per run |
|---|---|
| GPU time (L40S) | $0.00117 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3ร the typical price is reserved from your balance and settled afterwards.
- 1 credit = $0.01
Cost calculator
Price calculator
Typical according to the provider: about 1 s
Total
$0.12
12 credits
Per run
$0.0012 ยท 0.12 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
idefics3-8b- Developer
- Community
- Category
- Multimodal
- Input
- Text, Image
- Output
- Text
- Context window
- 8,192 tokens
- Max. output
- 4,096 tokens
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- September 23, 2026
Tags
- replicate
- multimodal
- vision-understanding
- huggingface
- open-weights
Frequently asked questions
What is Idefics3 8B?
Idefics3 8B is a model by Community in the Multimodal category.
How much does Idefics3 8B cost on Railwail?
On Railwail, Idefics3 8B costs โ $0.0012 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.
What is the context window of Idefics3 8B?
The context window of Idefics3 8B holds 8,192 tokens. One response can be up to 4,096 tokens long.
How fast is Idefics3 8B?
There are not enough measured runs of Idefics3 8B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Idefics3 8B better than BLIP?
That depends on the task. Idefics3 8B (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare Idefics3 8B and BLIPCan Idefics3 8B process images?
Yes. Idefics3 8B accepts images as input in addition to text.
Comparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.