Llama 3.2 90B Vision (multimodal)

MultimodalMevcut Değil
Meta tarafındanModel Kimliği: llama-3-2-90b-vision-mm

Meta's flagship vision-language model. 90B parameters, image understanding + chat, strong VQA performance.

Durum
Mevcut Değil
Bağlam
131.072 token
Maks. Çıkış
8.192 token
Giriş → Çıkış
Metin + Görüntü → Metin
Geliştirici
Meta
Güncellendi
25 Haziran 2026

Llama 3.2 90B Vision (multimodal) şu anda kullanılamıyor

Şu anda kullanılamıyor: bu model devre dışı bırakılmıştır.

Bu sayfadaki ayrıntıları yine de okuyabilirsiniz. Hemen karşılaştırılabilir bir modeli çalıştırmak için aşağıdaki mevcut alternatiflerden birini seçin.

Alternatiflere Git
01

Karşılaştırılabilir modeller

Bu kategorideki tümü
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ $0,00030/çalıştırma

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    $6,00/1M giriş

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    $3,60/1M giriş

02

Playground

Llama 3.2 90B Vision (multimodal)'ı deneyin

Sohbet

Şu anda kullanılamıyor

Şu anda kullanılamıyor: bu model devre dışı bırakılmıştır.

Oyun alanı devre dışıdır. Aynı kategoride karşılaştırılabilir modeller bulabilirsiniz: Alternatifleri görüntüle

Llama 3.2 90B Vision (multimodal)'ı deneyin

Bir mesaj gönderin. Yanıt model işini bitirdikten sonra tamamen gelir (akışsız).

Sistem istemi
Maks. yanıt uzunluğu (token)

Bu çalıştırma

Fiyat yok – şu anda kullanılamıyor.

Yeni misiniz?

Google ile kaydolduğunuzda 5 ücretsiz kredi ($0,05)

Kaydolduktan 24 saat sonra kullanılabilir, günde 5 çalıştırmaya kadar ve çalıştırma başına en fazla 2 kredi. Diğer oturum açma yöntemleri kredi olmadan başlar.

03

Llama 3.2 90B Vision (multimodal) Hakkında

Özet25 Haziran 2026 itibariyle

Llama 3.2 90B Vision (multimodal), Meta tarafından Multimodal kategorisinde geliştirilen bir modeldir. Llama 3.2 90B Vision (multimodal) şu anda Railwail üzerinde kullanılamıyor. Kontekst penceresi 131.072 token içerir ve bir yanıt en fazla 8.192 token uzunluğunda olabilir.

Arka plan

Meta AI (FAIR) hakkında

Kuruluş yılı 2013 · Menlo Park, California, USA

Meta AI is the research arm of Meta Platforms, established in 2013 as Facebook AI Research (FAIR) by Yann LeCun. FAIR has open-sourced many foundational models including PyTorch, RoBERTa, DETR, SAM and the LLaMA family. LLaMA 1 was released in February 2023, LLaMA 2 in July 2023, LLaMA 3 in April 2024 and LLaMA 3.1 (405B) in July 2024. Llama 3.2 launched in September 2024 at Meta Connect, introducing the first multimodal models in the LLaMA family (vision-enabled 11B and 90B) together with tiny on-device text-only siblings (1B, 3B). All Llama 3.2 vision weights are released under the Llama 3 Community Licence and are widely used by enterprise customers via Meta's partner ecosystem (Hugging Face, AWS Bedrock, Azure AI Studio, Google Vertex, Together AI, Groq, Fireworks).

Meta AI (FAIR) ziyaret edin

Mimari

Decoder-only Transformer with cross-attended vision encoder

Llama 3.2 90B Vision combines the 70B-parameter Llama 3.1 text backbone (extended to 90B with vision components) and a Vision Transformer image encoder integrated via cross-attention adapter layers, similar in spirit to Flamingo but reusing the LLaMA architecture. The vision tower processes each image to a sequence of visual tokens which are injected into specific cross-attention layers of the LLM decoder while the original text-only weights remain frozen during the multimodal training stage, preserving text-only performance. Pretraining used 6B image-text pairs followed by multi-stage supervised fine-tuning and Direct Preference Optimisation (DPO) on a curated set of image instructions, math and chart data. The model supports a 128K context window and accepts up to 1120x1120 image inputs natively (with tiling for larger images). It does not support video or audio. Llama 3.2 90B Vision is released under the Llama 3 Community Licence (free for commercial use under 700M MAU).

Parametreler
90B
Bağlam
128.000 token

Yetenekler

  • Open-weights 90B vision-language model under Llama 3 Community Licence
  • 128K token context window
  • Image input up to 1120x1120 with tiling for larger images
  • Chart, diagram, OCR and document understanding
  • Strong on MMMU, MathVista, ChartQA and DocVQA among open-weights models
  • Multilingual: English, German, French, Italian, Portuguese, Spanish, Hindi, Thai
  • Tool use and JSON output via Llama 3.1 alignment recipe
  • Best for: open-weights multimodal apps, on-premise document AI, indie research

Eğitim ve lisans

Pretrained on 6B image-text pairs from public web and licensed sources; supervised fine-tuning and DPO on curated multimodal instruction data. Text knowledge inherited from Llama 3.1 (15T tokens).

Lisans: Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.

Güvenlik testleri: Meta publishes a comprehensive model card with red-team findings on CBRN, child-safety and hate-speech vectors, plus Llama Guard 3 and Prompt Guard 2 companion models for production safety.

Bilinen sınırlamalar

  • No video or audio input
  • Latency and cost dominated by 90B params; requires multi-GPU serving
  • Licence restricts the largest hyperscaler use cases
  • Vision quality below GPT-4o and Claude 3.5 Sonnet on hardest charts
  • English-centric in vision domain
04

Fiyatlandırma

Şu anda kullanılamıyor: bu model devre dışı bırakılmıştır. Bu modelin şu anda bir fiyatı yok, bu nedenle çalıştırılamıyor.

05

API

Llama 3.2 90B Vision (multimodal) öğesini Railwail API anahtarınızla çağırın. İstekte bu model kimliğini kullanın:
llama-3-2-90b-vision-mmAPI BelgeleriAPI Anahtarı Al

Şu anda kullanılamıyor

Modelin doğrulanmış bir fiyatı yok veya devre dışı bırakılmıştır; API çağrıları reddedilir.

06

Özellikler

Model Kimliği
llama-3-2-90b-vision-mm
Geliştirici
Meta
Kategori
Multimodal
Giriş
Metin, Görüntü
Çıkış
Metin
Bağlam penceresi
131.072 token
Maks. çıkış
8.192 token
Model boyutu
90B
Lisans
Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.
Katalog girişi güncellendi
25 Haziran 2026

Etiketler

  • meta
  • llama
  • multimodal
  • vision
  • open-weights
07

Kullanım Alanları

Nerelerde Kullanılır

  • Open-weights document AI and OCR pipelines
  • On-premise vision-language assistants
  • Chart and diagram understanding for analytics
  • Compliance and regulated-industry multimodal apps
  • Research baselines for vision-language models
08

Sık Sorulan Sorular

Llama 3.2 90B Vision (multimodal) nedir?

Llama 3.2 90B Vision (multimodal), Meta tarafından Multimodal kategorisinde oluşturulan bir modeldir. Railwail'de listelenmiştir ancak şu anda çalıştırılamaz.

Llama 3.2 90B Vision (multimodal) Railwail'de ne kadar maliyetlidir?

Llama 3.2 90B Vision (multimodal) şu anda Railwail'de çalıştırılamaz, bu nedenle mevcut bir fiyat yoktur. Fiyatları olan mevcut alternatifler bu sayfanın aşağısında listelenmiştir.

Llama 3.2 90B Vision (multimodal)'nin kontekst penceresi ne kadardır?

Llama 3.2 90B Vision (multimodal)'nin kontekst penceresi 131.072 token içerir. Bir yanıt en fazla 8.192 token uzunluğunda olabilir.

Llama 3.2 90B Vision (multimodal) ne kadar hızlıdır?

Railwail'de Llama 3.2 90B Vision (multimodal) için henüz çalıştırma süresi belirtmek için yeterli ölçülen çalıştırma yoktur. Bu, giriş, ayarlar ve sağlayıcıdaki yüke bağlıdır.

Llama 3.2 90B Vision (multimodal), BLIP'dan daha iyi midir?

Bu göreve bağlıdır. Llama 3.2 90B Vision (multimodal) (Meta) ve BLIP (Salesforce) her ikisi de Multimodal kategorisinde modellerdir. Karşılaştırma sayfası fiyatlarını ve özelliklerini yan yana gösterir.

Llama 3.2 90B Vision (multimodal) ve BLIP'ı karşılaştır

Llama 3.2 90B Vision (multimodal) görüntüleri işleyebilir mi?

Evet. Llama 3.2 90B Vision (multimodal) metne ek olarak giriş olarak görüntüleri kabul eder.

Llama 3.2 90B Vision (multimodal)'yi şu anda kullanabilir miyim?

Şu anda kullanılamıyor: bu model devre dışı bırakılmıştır. Sayfa çevrimiçi kalır; aynı kategoriden mevcut alternatifler aşağıda listelenmiştir.

Tüm Modeller Tek Bir API Aracılığıyla

Railwail'deki her model için bir API anahtarı. Kullanım, ön ödemeli kredilerden tahsil edilir, 1 kredi = $0,01.