Llama 3.2 90B Vision (multimodal)

MultimodaalNiet beschikbaar
van MetaModel-ID: llama-3-2-90b-vision-mm

Meta's flagship vision-language model. 90B parameters, image understanding + chat, strong VQA performance.

Status
Niet beschikbaar
Context
131.072 tokens
Max. uitvoer
8.192 tokens
Invoer → Uitvoer
Tekst + Afbeelding → Tekst
Ontwikkelaar
Meta
Bijgewerkt
25 juni 2026

Llama 3.2 90B Vision (multimodal) is momenteel niet beschikbaar

Momenteel niet beschikbaar: dit model is gedeactiveerd.

Je kunt de details op deze pagina nog steeds lezen. Kies een van de beschikbare alternatieven hieronder om direct een vergelijkbaar model uit te voeren.

Naar alternatieven
01

Vergelijkbare modellen

Alle in deze categorie
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ US$ 0,00030/uitvoering

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    US$ 6,00/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    US$ 3,60/1M in

02

Playground

Llama 3.2 90B Vision (multimodal) proberen

Chat

Momenteel niet beschikbaar

Momenteel niet beschikbaar: dit model is gedeactiveerd.

De playground is uitgeschakeld. Vergelijkbare modellen vind je in dezelfde categorie: Alternatieven bekijken

Llama 3.2 90B Vision (multimodal) proberen

Stuur een bericht. Het antwoord komt volledig binnen zodra het model klaar is (geen streaming).

Systeemprompt
Max. antwoordlengte (tokens)

Deze uitvoering

Geen prijs – momenteel niet beschikbaar.

Nieuw hier?

5 gratis credits (US$ 0,05) wanneer je je aanmeldt met Google

Bruikbaar 24 uur na aanmelding, tot 5 uitvoeringen per dag en maximaal 2 credits per uitvoering. Andere aanmeldmethoden starten zonder credits.

03

Over Llama 3.2 90B Vision (multimodal)

SamengevatPer 25 juni 2026

Llama 3.2 90B Vision (multimodal) is een model van Meta in de categorie Multimodaal. Llama 3.2 90B Vision (multimodal) is momenteel niet beschikbaar op Railwail. Het contextvenster bevat 131.072 tokens, en een antwoord kan tot 8.192 tokens lang zijn.

Achtergrond

Over Meta AI (FAIR)

Opgericht 2013 · Menlo Park, California, USA

Meta AI is the research arm of Meta Platforms, established in 2013 as Facebook AI Research (FAIR) by Yann LeCun. FAIR has open-sourced many foundational models including PyTorch, RoBERTa, DETR, SAM and the LLaMA family. LLaMA 1 was released in February 2023, LLaMA 2 in July 2023, LLaMA 3 in April 2024 and LLaMA 3.1 (405B) in July 2024. Llama 3.2 launched in September 2024 at Meta Connect, introducing the first multimodal models in the LLaMA family (vision-enabled 11B and 90B) together with tiny on-device text-only siblings (1B, 3B). All Llama 3.2 vision weights are released under the Llama 3 Community Licence and are widely used by enterprise customers via Meta's partner ecosystem (Hugging Face, AWS Bedrock, Azure AI Studio, Google Vertex, Together AI, Groq, Fireworks).

Meta AI (FAIR) bezoeken

Architectuur

Decoder-only Transformer with cross-attended vision encoder

Llama 3.2 90B Vision combines the 70B-parameter Llama 3.1 text backbone (extended to 90B with vision components) and a Vision Transformer image encoder integrated via cross-attention adapter layers, similar in spirit to Flamingo but reusing the LLaMA architecture. The vision tower processes each image to a sequence of visual tokens which are injected into specific cross-attention layers of the LLM decoder while the original text-only weights remain frozen during the multimodal training stage, preserving text-only performance. Pretraining used 6B image-text pairs followed by multi-stage supervised fine-tuning and Direct Preference Optimisation (DPO) on a curated set of image instructions, math and chart data. The model supports a 128K context window and accepts up to 1120x1120 image inputs natively (with tiling for larger images). It does not support video or audio. Llama 3.2 90B Vision is released under the Llama 3 Community Licence (free for commercial use under 700M MAU).

Parameters
90B
Context
128.000 tokens

Mogelijkheden

  • Open-weights 90B vision-language model under Llama 3 Community Licence
  • 128K token context window
  • Image input up to 1120x1120 with tiling for larger images
  • Chart, diagram, OCR and document understanding
  • Strong on MMMU, MathVista, ChartQA and DocVQA among open-weights models
  • Multilingual: English, German, French, Italian, Portuguese, Spanish, Hindi, Thai
  • Tool use and JSON output via Llama 3.1 alignment recipe
  • Best for: open-weights multimodal apps, on-premise document AI, indie research

Training & licentie

Pretrained on 6B image-text pairs from public web and licensed sources; supervised fine-tuning and DPO on curated multimodal instruction data. Text knowledge inherited from Llama 3.1 (15T tokens).

Licentie: Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.

Veiligheidstests: Meta publishes a comprehensive model card with red-team findings on CBRN, child-safety and hate-speech vectors, plus Llama Guard 3 and Prompt Guard 2 companion models for production safety.

Bekende beperkingen

  • No video or audio input
  • Latency and cost dominated by 90B params; requires multi-GPU serving
  • Licence restricts the largest hyperscaler use cases
  • Vision quality below GPT-4o and Claude 3.5 Sonnet on hardest charts
  • English-centric in vision domain
04

Prijzen

Momenteel niet beschikbaar: dit model is gedeactiveerd. Er is momenteel geen prijs voor dit model, dus het kan niet worden uitgevoerd.

05

API

Roep Llama 3.2 90B Vision (multimodal) aan met je Railwail API-sleutel. Gebruik deze model-ID in het verzoek:

Momenteel niet beschikbaar

Het model heeft geen geverifieerde prijs of is gedeactiveerd; API-aanroepen worden geweigerd.

06

Specificaties

Model-ID
llama-3-2-90b-vision-mm
Ontwikkelaar
Meta
Categorie
Multimodaal
Invoer
Tekst, Afbeelding
Uitvoer
Tekst
Contextvenster
131.072 tokens
Max. uitvoer
8.192 tokens
Modelgrootte
90B
Licentie
Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.
Catalogusitem bijgewerkt
25 juni 2026

Tags

  • meta
  • llama
  • multimodal
  • vision
  • open-weights
07

Gebruiksscenario's

Waarvoor het wordt gebruikt

  • Open-weights document AI and OCR pipelines
  • On-premise vision-language assistants
  • Chart and diagram understanding for analytics
  • Compliance and regulated-industry multimodal apps
  • Research baselines for vision-language models
08

Veelgestelde vragen

Wat is Llama 3.2 90B Vision (multimodal)?

Llama 3.2 90B Vision (multimodal) is een model van Meta in de categorie Multimodaal. Het staat in de Railwail-catalogus, maar kan momenteel niet worden uitgevoerd.

Hoeveel kost Llama 3.2 90B Vision (multimodal) op Railwail?

Llama 3.2 90B Vision (multimodal) kan momenteel niet op Railwail worden uitgevoerd, dus er is geen huidige prijs. Beschikbare alternatieven met prijzen staan verderop op deze pagina.

Wat is het contextvenster van Llama 3.2 90B Vision (multimodal)?

Het contextvenster van Llama 3.2 90B Vision (multimodal) bevat 131.072 tokens. Een antwoord kan tot 8.192 tokens lang zijn.

Hoe snel is Llama 3.2 90B Vision (multimodal)?

Er zijn nog niet genoeg gemeten runs van Llama 3.2 90B Vision (multimodal) op Railwail om een uitvoeringstijd op te geven. Dit hangt af van de invoer, de instellingen en de belasting bij de provider.

Is Llama 3.2 90B Vision (multimodal) beter dan BLIP?

Dat hangt van de taak af. Llama 3.2 90B Vision (multimodal) (Meta) en BLIP (Salesforce) zijn beide modellen in de categorie Multimodaal. De vergelijkingspagina toont hun prijzen en specificaties naast elkaar.

Llama 3.2 90B Vision (multimodal) en BLIP vergelijken

Kan Llama 3.2 90B Vision (multimodal) afbeeldingen verwerken?

Ja. Llama 3.2 90B Vision (multimodal) accepteert afbeeldingen als invoer naast tekst.

Kan ik Llama 3.2 90B Vision (multimodal) nu gebruiken?

Momenteel niet beschikbaar: dit model is gedeactiveerd. De pagina blijft online; beschikbare alternatieven uit dezelfde categorie staan verderop.

Alle modellen via één API

Één API-sleutel voor elk model op Railwail. Gebruik wordt afgerekend via vooraf gekochte credits, 1 credit = US$ 0,01.