Llama 3.2 Vision 90B

MultimodalAvailable
by CommunityModel ID: llama-3-2-vision-90b

Meta Llama 3.2 90B Vision. Largest open-weights Llama vision model. Strong visual reasoning, chart, OCR and document understanding.

Price
โ‰ˆ US$0.0070/run
Context
131,072 tokens
Max. output
4,096 tokens
Input โ†’ output
Text + Image โ†’ Text
Developer
Community
Updated
23 September 2026
01

Playground

Try Llama 3.2 Vision 90B

No input form

โ‰ˆ US$0.0070/run

No input form for this model yet

Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.

02

Examples

Real outputs from the public examples of this model on Replicate, with the prompt and settings that produced them. They were not generated live on this page.
  • InputInput

    Prompt

    Which era does this piece belong to? Give details about the era.

    Response

    The image features a ceiling painting or mural, which is characteristic of the Rococo style. This art movement emerged in Europe during the 18th century, roughly between 1720 and 1770. **Key Characteristics:** * Lighter, more delicate colors compared to Baroque art * Use of pastel shades and soft hues **Associated Periods:** * Late Baroque period (1685-1750) * Early Classical period (1750-1820) These periods are marked by a shift away from the grandeur and drama of Baroque art towards a more refined and elegant aesthetic.

  • InputInput

    Prompt

    Where to find this monument? Can you give me other recommendations around the area?

    Response

    The monument in the image is located at Wat Phra Kaew, also known as the Temple of the Emerald Buddha, which is a complex of buildings within the Grand Palace in Bangkok, Thailand. The temple is situated on Na Phra Lan Road in Phra Nakhon District, Bangkok. **Other Recommendations Around the Area:** * **The Grand Palace**: A must-visit attraction in Bangkok, the palace was the official residence of the king and queen of Thailand until 1925. * **Wat Arun (Temple of Dawn)**: Located on the west bank of the Chao Phraya River, this temple is known for its stunning architecture and beautiful views of the river. * **Wat Pho (Temple of the Reclining Buddha)**: Home to a massive reclining Buddha statue, Wat Pho is one of the largest temples in Bangkok and offers traditional Thai massage services. * **Boat Ride along the Chao Phraya River**: Take a boat ride along the river to see the city from a different perspective and visit other attractions such as Wat Arun and Wat Pho. * **Floating Markets**: Visit one of the many floating markets in Bangkok, such as Damnoen Saduak Floating Market or Taling Chan Floating Market, to experience traditional Thai culture and try local food. * **Street Food**: Try some of the delicious street food available throughout Bangkok, including dishes like pad thai, tom yum soup, and mango sticky rice. * **MBK Center**: A large shopping mall with over 2,000 stores selling everything from clothing to electronics to souvenirs. * **Siam Paragon**: Another poโ€ฆ

  • InputInput

    Prompt

    What happens when you take out white cat from this chain?

    Response

    **Food Web Analysis: Impact of Removing the Wild Cat** To understand the effects of removing the wild cat from this food web, we need to examine its role and connections within the ecosystem. **Position in the Food Web** * The wild cat is a predator that feeds on rabbits. * Rabbits are preyed upon by owls but also consume green plants. **Impact on Prey Population (Rabbits)** * Removing the wild cat would reduce predation pressure on rabbits. * This could lead to an increase in the rabbit population, as one of their predators is removed. **Potential Effects on Other Species** * An increased rabbit population could result in: + Overgrazing: More rabbits consuming green plants could lead to a decrease in plant biomass and diversity. + Impact on Owl Population: With fewer wild cats competing for rabbit prey, the owl population might increase due to the abundance of rabbits. **Conclusion** Removing the wild cat from this food web would likely cause an increase in the rabbit population. This, in turn, could lead to overgrazing by rabbits and potentially impact the owl population positively due to increased availability of their primary food source.

03

About Llama 3.2 Vision 90B

TL;DRAs of 23 September 2026

Llama 3.2 Vision 90B is a model by Community in the Multimodal category. On Railwail, Llama 3.2 Vision 90B costs โ‰ˆ US$0.0070 per run. The context window holds 131,072 tokens, and one response can be up to 4,096 tokens long.

04

Pricing

Prices in US dollars. Usage is charged from prepaid credits.
Typical run (โ‰ˆ 4 s on A100 (80GB))US$0.0070 per run
GPU time (A100 (80GB))US$0.00168 per GPU second
  • Billed by the GPU time the run actually takes. When the run starts, 3ร— the typical price is reserved from your balance and settled afterwards.
  • 1 credit = US$0.01

Cost calculator

Price calculator

s

Typical according to the provider: about 4.1 s

Total

US$0.70

70 credits

Per run

US$0.007 ยท 0.7 credits

Billed by the actual GPU time; this is an estimate.

05

API

Call Llama 3.2 Vision 90B with your Railwail API key. Use this model ID in the request:

No verified API example

The public API passes a different input format than this model needs. Use the playground above.

06

Specifications

Model ID
llama-3-2-vision-90b
Developer
Community
Category
Multimodal
Input
Text, Image
Output
Text
Context window
131,072 tokens
Max. output
4,096 tokens
Billing
By usage (tokens or GPU time)
Catalog entry updated
23 September 2026

Tags

  • replicate
  • multimodal
  • vision-understanding
  • meta
  • open-weights
07

Frequently asked questions

What is Llama 3.2 Vision 90B?

Llama 3.2 Vision 90B is a model by Community in the Multimodal category.

How much does Llama 3.2 Vision 90B cost on Railwail?

On Railwail, Llama 3.2 Vision 90B costs โ‰ˆ US$0.0070 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals US$0.01.

What is the context window of Llama 3.2 Vision 90B?

The context window of Llama 3.2 Vision 90B holds 131,072 tokens. One response can be up to 4,096 tokens long.

How fast is Llama 3.2 Vision 90B?

There are not enough measured runs of Llama 3.2 Vision 90B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.

Is Llama 3.2 Vision 90B better than BLIP?

That depends on the task. Llama 3.2 Vision 90B (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.

Compare Llama 3.2 Vision 90B and BLIP

Can Llama 3.2 Vision 90B process images?

Yes. Llama 3.2 Vision 90B accepts images as input in addition to text.

08

Comparable models

All in this category
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    โ‰ˆ US$0.00030/run

    96 % cheaper per unit

    Compare Llama 3.2 Vision 90B vs. BLIP
  • pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.

    โ‰ˆ US$0.0457/run

    553 % more expensive per unit

    Compare Llama 3.2 Vision 90B vs. CLIP Interrogator
  • Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.

    โ‰ˆ US$0.0050/run

    29 % cheaper per unit

    Compare Llama 3.2 Vision 90B vs. Depth Anything v2

All models through one API

One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = US$0.01.