Llama 3.2 90B Vision (multimodal)

MultimodalIkke tilgjengelig
av MetaModell-ID: llama-3-2-90b-vision-mm

Meta's flagship vision-language model. 90B parameters, image understanding + chat, strong VQA performance.

Status
Ikke tilgjengelig
Kontekst
131 072 tokens
Maks. utdata
8 192 tokens
Input → output
Tekst + Bilde → Tekst
Utvikler
Meta
Oppdatert
25. juni 2026

Llama 3.2 90B Vision (multimodal) er for øyeblikket utilgjengelig

Ikke tilgjengelig for øyeblikket: dette modellen er deaktivert.

Du kan fortsatt lese detaljene på denne siden. Velg ett av de tilgjengelige alternativene nedenfor for å kjøre en sammenlignbar modell med en gang.

Gå til alternativer
01

Sammenlignbare modeller

Alle i denne kategorien
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 USD/kjøring

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 USD/1M inn

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 USD/1M inn

02

Playground

Prøv Llama 3.2 90B Vision (multimodal)

Chat

Ikke tilgjengelig for øyeblikket

Ikke tilgjengelig for øyeblikket: dette modellen er deaktivert.

Lekeplassen er deaktivert. Du finner sammenlignbare modeller i samme kategori: Se alternativer

Prøv Llama 3.2 90B Vision (multimodal)

Send en melding. Svaret kommer fullstendig når modellen er ferdig (ingen streaming).

Systemprompt
Maks. svarslengde (tokens)

Denne kjøringen

Ingen pris – ikke tilgjengelig for øyeblikket.

Ny her?

5 gratis credits (0,05 USD) når du registrerer deg med Google

Kan brukes 24 timer etter registrering, opptil 5 kjøringer per dag og maksimalt 2 credits per kjøring. Andre påloggingsmetoder starter uten credits.

03

Om Llama 3.2 90B Vision (multimodal)

Kort sagtPer 25. juni 2026

Llama 3.2 90B Vision (multimodal) er en modell fra Meta i kategorien Multimodal. Llama 3.2 90B Vision (multimodal) er for øyeblikket ikke tilgjengelig på Railwail. Kontekstvinduet inneholder 131 072 tokens, og ett svar kan være opptil 8 192 tokens langt.

Bakgrunn

Om Meta AI (FAIR)

Grunnlagt 2013 · Menlo Park, California, USA

Meta AI is the research arm of Meta Platforms, established in 2013 as Facebook AI Research (FAIR) by Yann LeCun. FAIR has open-sourced many foundational models including PyTorch, RoBERTa, DETR, SAM and the LLaMA family. LLaMA 1 was released in February 2023, LLaMA 2 in July 2023, LLaMA 3 in April 2024 and LLaMA 3.1 (405B) in July 2024. Llama 3.2 launched in September 2024 at Meta Connect, introducing the first multimodal models in the LLaMA family (vision-enabled 11B and 90B) together with tiny on-device text-only siblings (1B, 3B). All Llama 3.2 vision weights are released under the Llama 3 Community Licence and are widely used by enterprise customers via Meta's partner ecosystem (Hugging Face, AWS Bedrock, Azure AI Studio, Google Vertex, Together AI, Groq, Fireworks).

Besøk Meta AI (FAIR)

Arkitektur

Decoder-only Transformer with cross-attended vision encoder

Llama 3.2 90B Vision combines the 70B-parameter Llama 3.1 text backbone (extended to 90B with vision components) and a Vision Transformer image encoder integrated via cross-attention adapter layers, similar in spirit to Flamingo but reusing the LLaMA architecture. The vision tower processes each image to a sequence of visual tokens which are injected into specific cross-attention layers of the LLM decoder while the original text-only weights remain frozen during the multimodal training stage, preserving text-only performance. Pretraining used 6B image-text pairs followed by multi-stage supervised fine-tuning and Direct Preference Optimisation (DPO) on a curated set of image instructions, math and chart data. The model supports a 128K context window and accepts up to 1120x1120 image inputs natively (with tiling for larger images). It does not support video or audio. Llama 3.2 90B Vision is released under the Llama 3 Community Licence (free for commercial use under 700M MAU).

Parametere
90B
Kontekst
128 000 tokens

Evner

  • Open-weights 90B vision-language model under Llama 3 Community Licence
  • 128K token context window
  • Image input up to 1120x1120 with tiling for larger images
  • Chart, diagram, OCR and document understanding
  • Strong on MMMU, MathVista, ChartQA and DocVQA among open-weights models
  • Multilingual: English, German, French, Italian, Portuguese, Spanish, Hindi, Thai
  • Tool use and JSON output via Llama 3.1 alignment recipe
  • Best for: open-weights multimodal apps, on-premise document AI, indie research

Trening og lisens

Pretrained on 6B image-text pairs from public web and licensed sources; supervised fine-tuning and DPO on curated multimodal instruction data. Text knowledge inherited from Llama 3.1 (15T tokens).

Lisens: Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.

Sikkerhetstesting: Meta publishes a comprehensive model card with red-team findings on CBRN, child-safety and hate-speech vectors, plus Llama Guard 3 and Prompt Guard 2 companion models for production safety.

Kjente begrensninger

  • No video or audio input
  • Latency and cost dominated by 90B params; requires multi-GPU serving
  • Licence restricts the largest hyperscaler use cases
  • Vision quality below GPT-4o and Claude 3.5 Sonnet on hardest charts
  • English-centric in vision domain
04

Priser

Ikke tilgjengelig for øyeblikket: dette modellen er deaktivert. Det er ingen pris for denne modellen for øyeblikket, så den kan ikke kjøres.

05

API

Ring Llama 3.2 90B Vision (multimodal) med din Railwail API-nøkkel. Bruk denne modell-IDen i forespørselen:

Ikke tilgjengelig for øyeblikket

Modellen har ingen verifisert pris eller er deaktivert; API-kall blir avvist.

06

Spesifikasjoner

Modell-ID
llama-3-2-90b-vision-mm
Utvikler
Meta
Kategori
Multimodal
Inndata
Tekst, Bilde
Utdata
Tekst
Kontekstvindu
131 072 tokens
Maks. utdata
8 192 tokens
Modellstørrelse
90B
Lisens
Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.
Katalogoppføring oppdatert
25. juni 2026

Merkelapper

  • meta
  • llama
  • multimodal
  • vision
  • open-weights
07

Brukstilfeller

Hva det brukes til

  • Open-weights document AI and OCR pipelines
  • On-premise vision-language assistants
  • Chart and diagram understanding for analytics
  • Compliance and regulated-industry multimodal apps
  • Research baselines for vision-language models
08

Ofte stilte spørsmål

Hva er Llama 3.2 90B Vision (multimodal)?

Llama 3.2 90B Vision (multimodal) er en modell fra Meta i kategorien Multimodal. Den er oppført på Railwail, men kan ikke kjøres for øyeblikket.

Hvor mye koster Llama 3.2 90B Vision (multimodal) på Railwail?

Llama 3.2 90B Vision (multimodal) kan ikke kjøres på Railwail for øyeblikket, så det finnes ingen gjeldende pris. Tilgjengelige alternativer med priser er oppført lenger ned på denne siden.

Hvor stort er kontekstvinduet til Llama 3.2 90B Vision (multimodal)?

Kontekstvinduet til Llama 3.2 90B Vision (multimodal) inneholder 131 072 tokens. Ett svar kan være opptil 8 192 tokens langt.

Hvor rask er Llama 3.2 90B Vision (multimodal)?

Det finnes ennå ikke nok målte kjøringer av Llama 3.2 90B Vision (multimodal) på Railwail til å angi en kjøretid. Det avhenger av inndataene, innstillingene og belastningen hos leverandøren.

Er Llama 3.2 90B Vision (multimodal) bedre enn BLIP?

Det avhenger av oppgaven. Llama 3.2 90B Vision (multimodal) (Meta) og BLIP (Salesforce) er begge modeller i kategorien Multimodal. Sammenligningssiden viser prisene og spesifikasjonene deres side ved side.

Sammenlign Llama 3.2 90B Vision (multimodal) og BLIP

Kan Llama 3.2 90B Vision (multimodal) behandle bilder?

Ja. Llama 3.2 90B Vision (multimodal) godtar bilder som inndata i tillegg til tekst.

Kan jeg bruke Llama 3.2 90B Vision (multimodal) akkurat nå?

Ikke tilgjengelig for øyeblikket: dette modellen er deaktivert. Siden forblir online; tilgjengelige alternativer fra samme kategori er oppført lenger ned.

Alle modeller gjennom én API

Én API-nøkkel for alle modeller på Railwail. Bruk belastes fra forhåndsbetalt kreditt, 1 kreditt = 0,01 USD.