Llama 3.2 Vision 11B (Ollama)
llama-3-2-vision-11b-ollamaMeta Llama 3.2 11B Vision served via Ollama on Replicate. Open-weights multimodal model for image captioning, document and chart reading, and visual question answering.
- Pris
- ≈ 0,0039 US$/kørsel
- Input → output
- Tekst + Billede → Tekst
- Udvikler
- Community
- Opdateret
- 23. september 2026
Playground
Prøv Llama 3.2 Vision 11B (Ollama)
Input & resultat
Denne kørsel
ca. 0,0039 US$ · 0,39 credits
0,0116Â US$ (1,16 credits) reserveres ved start; den faktiske GPU-tid faktureres.
Ny her?
10 gratis credits (0,10 US$) når du tilmelder dig med Google
Kan bruges 24 timer efter tilmelding, op til 5 kørsler pr. dag og højst 2 credits pr. kørsel. Andre login-metoder starter uden credits. Nok til 8 kørsler af denne model.
Examples
InputPrompt
Where to find this monument? Can you give me other recommendations around the area?
Response
**Wat Phra Kaew (Temple of the Emerald Buddha)** in Bangkok, Thailand. Located within the Grand Palace complex, Wat Phra Kaew is a highly revered and sacred site that houses the iconic Emerald Buddha image. The temple's design incorporates traditional Thai architecture with intricate details and ornate decorations, featuring an imposing spire topped by a golden finial adorned with images of mythical creatures. **Recommendations for Your Visit:** * **Grand Palace**: Explore the sprawling palace grounds, which include several other notable temples and buildings. * **Wat Arun (Temple of Dawn)**: A nearby temple famous for its distinctive spire covered in colorful ceramics and glints of gold. * **Chao Phraya River Cruise**: Take a boat ride along the river to admire the city's skyline and learn about its history. * **Jim Thompson House Museum**: Visit this museum dedicated to Thai art and culture, featuring traditional architecture and beautiful gardens. Wat Phra Kaew is an absolute must-visit attraction when in Bangkok. With these recommendations, you'll have a well-rounded experience of the city's rich cultural heritage.
InputPrompt
How long does it take from invoice date to due date? Be short and concise.
Response
From the invoice date of 11/02/2019 to the due date of 26/02/2019, there is a difference of **15 days**. This calculation was performed by subtracting the invoice date from the due date: 26/02/2019 - 11/02/2019 = 15
InputPrompt
What happens when you take out white cat from this chain?
Response
The white cat is not present in the food web. The correct term for a wild animal would be "wildcat." If we were to remove a wildcat from the ecosystem, it would no longer be able to hunt or prey on other animals. This could have cascading effects throughout the entire ecosystem. *Answer*: Wildcat will lose its source of food and nutrients
Om Llama 3.2 Vision 11B (Ollama)
Llama 3.2 Vision 11B (Ollama) er en model af Community i kategorien Multimodal. På Railwail koster Llama 3.2 Vision 11B (Ollama) ≈ 0,0039 US$ pr. kørsel.
Priser
| Typisk kørsel (≈ 3 s på L40S) | 0,0039 US$ pr. kørsel |
|---|---|
| GPU-tid (L40S) | 0,00117Â US$ pr. GPU-sekund |
- Faktureres efter den GPU-tid, som kørslen faktisk tager. Når kørslen starter, reserveres 3× den typiske pris fra din saldo og afregnes efterfølgende.
- 1 kredit = 0,01Â US$
Omkostningsberegner
Prisberegner
Typisk ifølge udbyderen: ca. 3,3 s
I alt
0,39Â US$
39 credits
Pr. kørsel
0,0039 US$ · 0,39 credits
Fakturering efter faktisk GPU-tid; dette er et estimat.
API
Intet bekræftet API-eksempel
Den offentlige API sender et andet inputformat end det, som denne model har brug for. Brug playground ovenfor.
Specifikationer
- Model-ID
llama-3-2-vision-11b-ollama- Udvikler
- Community
- Kategori
- Multimodal
- Input
- Tekst, Billede
- Output
- Tekst
- Fakturering
- Efter forbrug (tokens eller GPU-tid)
- Katalogelement opdateret
- 23. september 2026
Inputparametre
Inputs og indstillinger fra modellens inputskema. Eksemplet i API-afsnittet viser, hvilke af dem API'en accepterer.
promptpåkrævetQuestion about the image
Type: TekstStandard: –Tilladte værdier: op til 16.000 tegnimage_urlImage URL to analyze
Type: TekstStandard: –Tilladte værdier: –max_tokensType: HeltalStandard:1024Tilladte værdier: 1 til 4.096temperatureType: TalStandard:0.7Tilladte værdier: 0 til 2
Tags
- replicate
- meta
- llama
- vision-understanding
- open-weights
- ollama
Ofte stillede spørgsmål
Hvad er Llama 3.2 Vision 11B (Ollama)?
Llama 3.2 Vision 11B (Ollama) er en model fra Community i kategorien Multimodal.
Hvad koster Llama 3.2 Vision 11B (Ollama) på Railwail?
På Railwail koster Llama 3.2 Vision 11B (Ollama) ≈ 0,0039 US$ pr. kørsel. Du betaler for det, som hver anmodning faktisk bruger. Forbrug betales fra forudbetalte credits; 1 credit svarer til 0,01 US$.
Hvilke indstillinger understøtter Llama 3.2 Vision 11B (Ollama)?
Ifølge dens inputskema kender Llama 3.2 Vision 11B (Ollama) disse parametre: prompt (op til 16.000 tegn), image_url, max_tokens (1 til 4.096) og temperature (0 til 2).
Hvor hurtig er Llama 3.2 Vision 11B (Ollama)?
Der er endnu ikke nok målte kørsler af Llama 3.2 Vision 11B (Ollama) på Railwail til at angive en udførelsestid. Det afhænger af inputtet, indstillingerne og belastningen hos provideren.
Er Llama 3.2 Vision 11B (Ollama) bedre end BLIP?
Det afhænger af opgaven. Llama 3.2 Vision 11B (Ollama) (Community) og BLIP (Salesforce) er begge modeller i kategorien Multimodal. Sammenligningssiden viser deres priser og specifikationer side om side.
Sammenlign Llama 3.2 Vision 11B (Ollama) og BLIPKan Llama 3.2 Vision 11B (Ollama) behandle billeder?
Ja. Llama 3.2 Vision 11B (Ollama) accepterer billeder som input ud over tekst.
Sammenlignelige modeller
Alle i denne kategori- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
≈ 0,0457 US$/kørsel
1.072 % dyrere pr. enhed
Sammenlign Llama 3.2 Vision 11B (Ollama) og CLIP Interrogator - Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
≈ 0,0050 US$/kørsel
28 % dyrere pr. enhed
Sammenlign Llama 3.2 Vision 11B (Ollama) og Depth Anything v2
Alle modeller via én API
En API-nøgle til alle modeller på Railwail. Forbrug debiteres fra forudbetalte credits, 1 credit = 0,01 US$.