Llama 3.2 Vision 11B (Ollama)

MultimodalAvailable
by CommunityModel ID: llama-3-2-vision-11b-ollama

Meta Llama 3.2 11B Vision served via Ollama on Replicate. Open-weights multimodal model for image captioning, document and chart reading, and visual question answering.

Price
β‰ˆ $0.0039/run
Input β†’ output
Text + Image β†’ Text
Developer
Community
Updated
September 23, 2026
01

Playground

Try Llama 3.2 Vision 11B (Ollama)

Input & output

β‰ˆ $0.0039/run
Try Llama 3.2 Vision 11B (Ollama)

0 / 16,000

Question about the image

Image URL to analyze

Advanced settings (2)
Output
The answer appears here.

This run

about $0.0039 Β· 0.39 credits

$0.0116 (1.16 credits) are reserved at the start; the actual GPU time is billed.

New here?

10 free credits ($0.10) when you sign up with Google

Usable 24 hours after sign-up, up to 5 runs per day and at most 2 credits per run. Other sign-in methods start without credits. Enough for 8 runs of this model.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • InputInput

    Prompt

    Where to find this monument? Can you give me other recommendations around the area?

    Response

    **Wat Phra Kaew (Temple of the Emerald Buddha)** in Bangkok, Thailand. Located within the Grand Palace complex, Wat Phra Kaew is a highly revered and sacred site that houses the iconic Emerald Buddha image. The temple's design incorporates traditional Thai architecture with intricate details and ornate decorations, featuring an imposing spire topped by a golden finial adorned with images of mythical creatures. **Recommendations for Your Visit:** * **Grand Palace**: Explore the sprawling palace grounds, which include several other notable temples and buildings. * **Wat Arun (Temple of Dawn)**: A nearby temple famous for its distinctive spire covered in colorful ceramics and glints of gold. * **Chao Phraya River Cruise**: Take a boat ride along the river to admire the city's skyline and learn about its history. * **Jim Thompson House Museum**: Visit this museum dedicated to Thai art and culture, featuring traditional architecture and beautiful gardens. Wat Phra Kaew is an absolute must-visit attraction when in Bangkok. With these recommendations, you'll have a well-rounded experience of the city's rich cultural heritage.

  • InputInput

    Prompt

    How long does it take from invoice date to due date? Be short and concise.

    Response

    From the invoice date of 11/02/2019 to the due date of 26/02/2019, there is a difference of **15 days**. This calculation was performed by subtracting the invoice date from the due date: 26/02/2019 - 11/02/2019 = 15

  • InputInput

    Prompt

    What happens when you take out white cat from this chain?

    Response

    The white cat is not present in the food web. The correct term for a wild animal would be "wildcat." If we were to remove a wildcat from the ecosystem, it would no longer be able to hunt or prey on other animals. This could have cascading effects throughout the entire ecosystem. *Answer*: Wildcat will lose its source of food and nutrients

03

About Llama 3.2 Vision 11B (Ollama)

TL;DRAs of September 23, 2026

Llama 3.2 Vision 11B (Ollama) is a model by Community in the Multimodal category. On Railwail, Llama 3.2 Vision 11B (Ollama) costs β‰ˆ $0.0039 per run.

This endpoint runs Meta's Llama 3.2 11B Vision Instruct through Ollama on Replicate. It takes a single image plus a text prompt and answers questions, describes images, reads charts and documents and performs general visual reasoning. The 11B size keeps cost low while staying usable for caption generation, alt-text, receipt and form reading and screenshot understanding. Fully open weights from Meta.
04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (β‰ˆ 3 s on L40S)$0.0039 per run
GPU time (L40S)$0.00117 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3Γ— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = $0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 3.3 s

Total

$0.39

39 credits

Per run

$0.0039 Β· 0.39 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call Llama 3.2 Vision 11B (Ollama) with your Railwail API key. Use this model ID in the request:
llama-3-2-vision-11b-ollamaAPI documentationGet an API key

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
llama-3-2-vision-11b-ollama
Developer
Community
Category
Multimodal
Input
Text, Image
Output
Text
Billing
By usage (tokens or GPU time)
Catalog entry updated
September 23, 2026

Input parameters

Inputs and settings from the model's input schema. The example in the API section shows which of them the API accepts.

  • promptrequired

    Question about the image

    Type: Text
    Default: –
    Allowed values: up to 16,000 characters
  • image_url

    Image URL to analyze

    Type: Text
    Default: –
    Allowed values: –
  • max_tokens
    Type: Integer
    Default: 1024
    Allowed values: 1 to 4,096
  • temperature
    Type: Number
    Default: 0.7
    Allowed values: 0 to 2

Tags

  • replicate
  • meta
  • llama
  • vision-understanding
  • open-weights
  • ollama
07

Frequently asked questions

What is Llama 3.2 Vision 11B (Ollama)?

Llama 3.2 Vision 11B (Ollama) is a model by Community in the Multimodal category.

How much does Llama 3.2 Vision 11B (Ollama) cost on Railwail?

On Railwail, Llama 3.2 Vision 11B (Ollama) costs β‰ˆ $0.0039 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.

Which settings does Llama 3.2 Vision 11B (Ollama) support?

According to its input schema, Llama 3.2 Vision 11B (Ollama) knows these parameters: prompt (up to 16,000 characters), image_url, max_tokens (1 to 4,096), and temperature (0 to 2).

How fast is Llama 3.2 Vision 11B (Ollama)?

There are not enough measured runs of Llama 3.2 Vision 11B (Ollama) on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Llama 3.2 Vision 11B (Ollama) better than BLIP?

That depends on the task. Llama 3.2 Vision 11B (Ollama) (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare Llama 3.2 Vision 11B (Ollama) and BLIP

Can Llama 3.2 Vision 11B (Ollama) process images?

Yes. Llama 3.2 Vision 11B (Ollama) accepts images as input in addition to text.

08

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    β‰ˆ $0.00030/run

    92 % cheaper per unit

    Compare Llama 3.2 Vision 11B (Ollama) vs. BLIP
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    β‰ˆ $0.0457/run

    1,072 % more expensive per unit

    Compare Llama 3.2 Vision 11B (Ollama) vs. CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    β‰ˆ $0.0050/run

    28 % more expensive per unit

    Compare Llama 3.2 Vision 11B (Ollama) vs. Depth Anything v2

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.