Llama 3.2 Vision 90B
llama-3-2-vision-90bMeta Llama 3.2 90B Vision. Largest open-weights Llama vision model. Strong visual reasoning, chart, OCR and document understanding.
- Price
- โ $0.0070/run
- Context
- 131,072 tokens
- Max. output
- 4,096 tokens
- Input โ output
- Text + Image โ Text
- Developer
- Community
- Updated
- September 23, 2026
Playground
Try Llama 3.2 Vision 90B
No input form
No input form for this model yet
Its inputs are not documented yet. So that no run fails on a wrong input, we don't offer a form here. Pick a comparable model instead.
Examples
InputPrompt
Which era does this piece belong to? Give details about the era.
Response
The image features a ceiling painting or mural, which is characteristic of the Rococo style. This art movement emerged in Europe during the 18th century, roughly between 1720 and 1770. **Key Characteristics:** * Lighter, more delicate colors compared to Baroque art * Use of pastel shades and soft hues **Associated Periods:** * Late Baroque period (1685-1750) * Early Classical period (1750-1820) These periods are marked by a shift away from the grandeur and drama of Baroque art towards a more refined and elegant aesthetic.
InputPrompt
Where to find this monument? Can you give me other recommendations around the area?
Response
The monument in the image is located at Wat Phra Kaew, also known as the Temple of the Emerald Buddha, which is a complex of buildings within the Grand Palace in Bangkok, Thailand. The temple is situated on Na Phra Lan Road in Phra Nakhon District, Bangkok. **Other Recommendations Around the Area:** * **The Grand Palace**: A must-visit attraction in Bangkok, the palace was the official residence of the king and queen of Thailand until 1925. * **Wat Arun (Temple of Dawn)**: Located on the west bank of the Chao Phraya River, this temple is known for its stunning architecture and beautiful views of the river. * **Wat Pho (Temple of the Reclining Buddha)**: Home to a massive reclining Buddha statue, Wat Pho is one of the largest temples in Bangkok and offers traditional Thai massage services. * **Boat Ride along the Chao Phraya River**: Take a boat ride along the river to see the city from a different perspective and visit other attractions such as Wat Arun and Wat Pho. * **Floating Markets**: Visit one of the many floating markets in Bangkok, such as Damnoen Saduak Floating Market or Taling Chan Floating Market, to experience traditional Thai culture and try local food. * **Street Food**: Try some of the delicious street food available throughout Bangkok, including dishes like pad thai, tom yum soup, and mango sticky rice. * **MBK Center**: A large shopping mall with over 2,000 stores selling everything from clothing to electronics to souvenirs. * **Siam Paragon**: Another poโฆ
InputPrompt
What happens when you take out white cat from this chain?
Response
**Food Web Analysis: Impact of Removing the Wild Cat** To understand the effects of removing the wild cat from this food web, we need to examine its role and connections within the ecosystem. **Position in the Food Web** * The wild cat is a predator that feeds on rabbits. * Rabbits are preyed upon by owls but also consume green plants. **Impact on Prey Population (Rabbits)** * Removing the wild cat would reduce predation pressure on rabbits. * This could lead to an increase in the rabbit population, as one of their predators is removed. **Potential Effects on Other Species** * An increased rabbit population could result in: + Overgrazing: More rabbits consuming green plants could lead to a decrease in plant biomass and diversity. + Impact on Owl Population: With fewer wild cats competing for rabbit prey, the owl population might increase due to the abundance of rabbits. **Conclusion** Removing the wild cat from this food web would likely cause an increase in the rabbit population. This, in turn, could lead to overgrazing by rabbits and potentially impact the owl population positively due to increased availability of their primary food source.
About Llama 3.2 Vision 90B
Llama 3.2 Vision 90B is a model by Community in the Multimodal category. On Railwail, Llama 3.2 Vision 90B costs โ $0.0070 per run. The context window holds 131,072 tokens, and one response can be up to 4,096 tokens long.
Pricing
| Typical run (โ 4 s on A100 (80GB)) | $0.0070 per run |
|---|---|
| GPU time (A100 (80GB)) | $0.00168 per GPU second |
- Billed by the GPU time the run actually takes. When the run starts, 3ร the typical price is reserved from your balance and settled afterwards.
- 1 credit = $0.01
Cost calculator
Price calculator
Typical according to the provider: about 4.1 s
Total
$0.70
70 credits
Per run
$0.007 ยท 0.7 credits
Billed by the actual GPU time; this is an estimate.
API
No verified API example
The public API passes a different input format than this model needs. Use the playground above.
Specifications
- Model ID
llama-3-2-vision-90b- Developer
- Community
- Category
- Multimodal
- Input
- Text, Image
- Output
- Text
- Context window
- 131,072 tokens
- Max. output
- 4,096 tokens
- Billing
- By usage (tokens or GPU time)
- Catalog entry updated
- September 23, 2026
Tags
- replicate
- multimodal
- vision-understanding
- meta
- open-weights
Frequently asked questions
What is Llama 3.2 Vision 90B?
Llama 3.2 Vision 90B is a model by Community in the Multimodal category.
How much does Llama 3.2 Vision 90B cost on Railwail?
On Railwail, Llama 3.2 Vision 90B costs โ $0.0070 per run. You are charged for what each request actually uses. Usage is paid from prepaid credits; 1 credit equals $0.01.
What is the context window of Llama 3.2 Vision 90B?
The context window of Llama 3.2 Vision 90B holds 131,072 tokens. One response can be up to 4,096 tokens long.
How fast is Llama 3.2 Vision 90B?
There are not enough measured runs of Llama 3.2 Vision 90B on Railwail yet to state a run time. It depends on the input, the settings and the load at the provider.
Is Llama 3.2 Vision 90B better than BLIP?
That depends on the task. Llama 3.2 Vision 90B (Community) and BLIP (Salesforce) are both models in the Multimodal category. The comparison page shows their prices and specifications side by side.
Compare Llama 3.2 Vision 90B and BLIPCan Llama 3.2 Vision 90B process images?
Yes. Llama 3.2 Vision 90B accepts images as input in addition to text.
Comparable models
All in this category- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
- Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
All models through one API
One API key for every model on Railwail. Usage is charged from prepaid credits, 1 credit = $0.01.