Llama 3.2 Vision 11B (Ollama)
llama-3-2-vision-11b-ollamaMeta Llama 3.2 11B Vision served via Ollama on Replicate. Open-weights multimodal model for image captioning, document and chart reading, and visual question answering.
- Цена
- ≈ 0,0039 $/запуск
- Вход → выход
- Текст + Изображение → Текст
- Разработчик
- Community
- Обновлено
- 23 сентября 2026 г.
Playground
Попробовать Llama 3.2 Vision 11B (Ollama)
Входные данные и результат
Этот запуск
примерно 0,0039 $ · 0,39 кредитов
При запуске зарезервировано 0,0116 $ (1,16 кредитов); оплачивается фактическое время GPU.
Впервые здесь?
10 бесплатных кредитов (0,10 $) при регистрации через Google
Доступно 24 часов после регистрации, до 5 запусков в день и максимум 2 кредитов за запуск. Другие способы входа начинают без кредитов. Достаточно для 8 запусков этой модели.
Examples
InputPrompt
Where to find this monument? Can you give me other recommendations around the area?
Response
**Wat Phra Kaew (Temple of the Emerald Buddha)** in Bangkok, Thailand. Located within the Grand Palace complex, Wat Phra Kaew is a highly revered and sacred site that houses the iconic Emerald Buddha image. The temple's design incorporates traditional Thai architecture with intricate details and ornate decorations, featuring an imposing spire topped by a golden finial adorned with images of mythical creatures. **Recommendations for Your Visit:** * **Grand Palace**: Explore the sprawling palace grounds, which include several other notable temples and buildings. * **Wat Arun (Temple of Dawn)**: A nearby temple famous for its distinctive spire covered in colorful ceramics and glints of gold. * **Chao Phraya River Cruise**: Take a boat ride along the river to admire the city's skyline and learn about its history. * **Jim Thompson House Museum**: Visit this museum dedicated to Thai art and culture, featuring traditional architecture and beautiful gardens. Wat Phra Kaew is an absolute must-visit attraction when in Bangkok. With these recommendations, you'll have a well-rounded experience of the city's rich cultural heritage.
InputPrompt
How long does it take from invoice date to due date? Be short and concise.
Response
From the invoice date of 11/02/2019 to the due date of 26/02/2019, there is a difference of **15 days**. This calculation was performed by subtracting the invoice date from the due date: 26/02/2019 - 11/02/2019 = 15
InputPrompt
What happens when you take out white cat from this chain?
Response
The white cat is not present in the food web. The correct term for a wild animal would be "wildcat." If we were to remove a wildcat from the ecosystem, it would no longer be able to hunt or prey on other animals. This could have cascading effects throughout the entire ecosystem. *Answer*: Wildcat will lose its source of food and nutrients
О Llama 3.2 Vision 11B (Ollama)
Llama 3.2 Vision 11B (Ollama) — это модель от Community в категории Мультимодальные. На Railwail Llama 3.2 Vision 11B (Ollama) стоит ≈ 0,0039 $ за запуск.
Цены
| Типичный запуск (≈ 3 с на L40S) | 0,0039 $ за запуск |
|---|---|
| Время GPU (L40S) | 0,00117 $ за секунду GPU |
- Оплата рассчитывается по времени GPU, которое фактически требуется для запуска. При запуске из вашего баланса резервируется 3× типичная цена, которая затем списывается.
- 1 кредит = 0,01 $
Калькулятор стоимости
Калькулятор цен
По данным поставщика: примерно 3,3 с
Итого
0,39 $
39 кредитов
За запуск
0,0039 $ · 0,39 кредитов
Выставляется счет за фактическое время GPU; это оценка.
API
Нет проверенного примера API
Публичный API передает другой формат входных данных, чем требует эта модель. Используйте playground выше.
Спецификации
- ID модели
llama-3-2-vision-11b-ollama- Разработчик
- Community
- Категория
- Мультимодальные
- Входные данные
- Текст, Изображение
- Выходные данные
- Текст
- Тарификация
- По использованию (токены или время GPU)
- Запись в каталоге обновлена
- 23 сентября 2026 г.
Входные параметры
Входные данные и параметры из схемы входных данных модели. Пример в разделе API показывает, какие из них принимает API.
promptобязательноQuestion about the image
Тип: ТекстПо умолчанию: –Допустимые значения: до 16 000 символовimage_urlImage URL to analyze
Тип: ТекстПо умолчанию: –Допустимые значения: –max_tokensТип: Целое числоПо умолчанию:1024Допустимые значения: от 1 до 4 096temperatureТип: ЧислоПо умолчанию:0.7Допустимые значения: от 0 до 2
Теги
- replicate
- meta
- llama
- vision-understanding
- open-weights
- ollama
Часто задаваемые вопросы
Что такое Llama 3.2 Vision 11B (Ollama)?
Llama 3.2 Vision 11B (Ollama) — модель от Community в категории Мультимодальные.
Сколько стоит Llama 3.2 Vision 11B (Ollama) на Railwail?
На Railwail Llama 3.2 Vision 11B (Ollama) стоит ≈ 0,0039 $ за запуск. Вы платите за то, что фактически использует каждый запрос. Использование оплачивается предоплаченными кредитами; 1 кредит = 0,01 $.
Какие параметры поддерживает Llama 3.2 Vision 11B (Ollama)?
Согласно схеме входных данных, Llama 3.2 Vision 11B (Ollama) поддерживает эти параметры: prompt (до 16 000 символов), image_url, max_tokens (от 1 до 4 096) и temperature (от 0 до 2).
Насколько быстра Llama 3.2 Vision 11B (Ollama)?
На Railwail пока недостаточно измеренных запусков Llama 3.2 Vision 11B (Ollama), чтобы указать время выполнения. Оно зависит от входных данных, параметров и нагрузки на провайдера.
Llama 3.2 Vision 11B (Ollama) лучше, чем BLIP?
Это зависит от задачи. Llama 3.2 Vision 11B (Ollama) (Community) и BLIP (Salesforce) — обе модели в категории Мультимодальные. На странице сравнения показаны их цены и характеристики рядом.
Сравнить Llama 3.2 Vision 11B (Ollama) и BLIPМожет ли Llama 3.2 Vision 11B (Ollama) обрабатывать изображения?
Да. Llama 3.2 Vision 11B (Ollama) принимает изображения на входе в дополнение к тексту.
Сравнимые модели
Все в этой категории- BLIPSalesforce
Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.
- CLIP InterrogatorCommunity
pharmapsychotic's CLIP Interrogator. Takes an image and produces a Stable-Diffusion-style text prompt by combining BLIP captioning with CLIP to rank likely subjects, artists, mediums and styles. Commonly used to reverse-engineer a prompt from an existing picture.
≈ 0,0457 $/запуск
1 072 % дороже за единицу
Сравнить Llama 3.2 Vision 11B (Ollama) и CLIP Interrogator - Depth Anything v2Community
Monocular depth-estimation model trained on 595k labeled and 62M unlabeled images. Strong zero-shot generalization in indoor and outdoor scenes.
Все модели через один API
Один API-ключ для всех моделей на Railwail. Использование оплачивается из предоплаченных кредитов, 1 кредит = 0,01 $.