Llama 3.2 90B Vision (multimodal)

MultimodalIndisponibil
de MetaID model: llama-3-2-90b-vision-mm

Meta's flagship vision-language model. 90B parameters, image understanding + chat, strong VQA performance.

Status
Indisponibil
Context
131.072 tokeni
Ieșire max.
8.192 tokeni
Intrare → ieșire
Text + Imagine → Text
Dezvoltator
Meta
Actualizat
25 iunie 2026

Llama 3.2 90B Vision (multimodal) nu este disponibil în acest moment

Indisponibil în prezent: acest model a fost dezactivat.

Poți citi în continuare detaliile pe această pagină. Alege una dintre alternativele disponibile de mai jos pentru a rula imediat un model comparabil.

Mergi la alternative
01
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 USD/rulare

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 USD/1M in

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 USD/1M in

02

Playground

Încearcă Llama 3.2 90B Vision (multimodal)

Chat

Indisponibil în prezent

Indisponibil în prezent: acest model a fost dezactivat.

Playground-ul este dezactivat. Modele comparabile găsești în aceeași categorie: Vezi alternativele

Încearcă Llama 3.2 90B Vision (multimodal)

Trimite un mesaj. Răspunsul sosește complet când modelul termină (fără streaming).

Prompt de sistem
Lungimea max. a răspunsului (tokeni)

Această rulare

Fără preț – momentan indisponibil.

Nou aici?

5 credite gratuite (0,05 USD) când te înregistrezi cu Google

Utilizabil 24 ore după înregistrare, până la 5 rulări pe zi și maximum 2 credite pe rulare. Alte metode de conectare încep fără credite.

03

Despre Llama 3.2 90B Vision (multimodal)

Pe scurtDin 25 iunie 2026

Llama 3.2 90B Vision (multimodal) este un model de Meta din categoria Multimodal. Llama 3.2 90B Vision (multimodal) nu este disponibil în prezent pe Railwail. Fereastra de context conține 131.072 token-uri, iar un răspuns poate fi lung de până la 8.192 token-uri.

Fundal

Despre Meta AI (FAIR)

Fondat 2013 · Menlo Park, California, USA

Meta AI is the research arm of Meta Platforms, established in 2013 as Facebook AI Research (FAIR) by Yann LeCun. FAIR has open-sourced many foundational models including PyTorch, RoBERTa, DETR, SAM and the LLaMA family. LLaMA 1 was released in February 2023, LLaMA 2 in July 2023, LLaMA 3 in April 2024 and LLaMA 3.1 (405B) in July 2024. Llama 3.2 launched in September 2024 at Meta Connect, introducing the first multimodal models in the LLaMA family (vision-enabled 11B and 90B) together with tiny on-device text-only siblings (1B, 3B). All Llama 3.2 vision weights are released under the Llama 3 Community Licence and are widely used by enterprise customers via Meta's partner ecosystem (Hugging Face, AWS Bedrock, Azure AI Studio, Google Vertex, Together AI, Groq, Fireworks).

Vizitează Meta AI (FAIR)

Arhitectură

Decoder-only Transformer with cross-attended vision encoder

Llama 3.2 90B Vision combines the 70B-parameter Llama 3.1 text backbone (extended to 90B with vision components) and a Vision Transformer image encoder integrated via cross-attention adapter layers, similar in spirit to Flamingo but reusing the LLaMA architecture. The vision tower processes each image to a sequence of visual tokens which are injected into specific cross-attention layers of the LLM decoder while the original text-only weights remain frozen during the multimodal training stage, preserving text-only performance. Pretraining used 6B image-text pairs followed by multi-stage supervised fine-tuning and Direct Preference Optimisation (DPO) on a curated set of image instructions, math and chart data. The model supports a 128K context window and accepts up to 1120x1120 image inputs natively (with tiling for larger images). It does not support video or audio. Llama 3.2 90B Vision is released under the Llama 3 Community Licence (free for commercial use under 700M MAU).

Parametri
90B
Context
128.000 tokeni

Capabilități

  • Open-weights 90B vision-language model under Llama 3 Community Licence
  • 128K token context window
  • Image input up to 1120x1120 with tiling for larger images
  • Chart, diagram, OCR and document understanding
  • Strong on MMMU, MathVista, ChartQA and DocVQA among open-weights models
  • Multilingual: English, German, French, Italian, Portuguese, Spanish, Hindi, Thai
  • Tool use and JSON output via Llama 3.1 alignment recipe
  • Best for: open-weights multimodal apps, on-premise document AI, indie research

Antrenament & licență

Pretrained on 6B image-text pairs from public web and licensed sources; supervised fine-tuning and DPO on curated multimodal instruction data. Text knowledge inherited from Llama 3.1 (15T tokens).

Licență: Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.

Teste de siguranță: Meta publishes a comprehensive model card with red-team findings on CBRN, child-safety and hate-speech vectors, plus Llama Guard 3 and Prompt Guard 2 companion models for production safety.

Limitări cunoscute

  • No video or audio input
  • Latency and cost dominated by 90B params; requires multi-GPU serving
  • Licence restricts the largest hyperscaler use cases
  • Vision quality below GPT-4o and Claude 3.5 Sonnet on hardest charts
  • English-centric in vision domain
04

Prețuri

Indisponibil în prezent: acest model a fost dezactivat. Nu există preț pentru acest model în acest moment, deci nu poate fi executat.

05

API

Apelează Llama 3.2 90B Vision (multimodal) cu cheia ta API Railwail. Folosește acest ID de model în cerere:

Indisponibil în prezent

Modelul nu are un preț verificat sau este dezactivat; apelurile API sunt refuzate.

06

Specificații

ID model
llama-3-2-90b-vision-mm
Dezvoltator
Meta
Categorie
Multimodal
Intrare
Text, Imagine
Ieșire
Text
Fereastră de context
131.072 tokeni
Ieșire max.
8.192 tokeni
Dimensiune model
90B
Licență
Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.
Intrare catalog actualizată
25 iunie 2026

Etichete

  • meta
  • llama
  • multimodal
  • vision
  • open-weights
07

Cazuri de utilizare

Pentru ce se folosește

  • Open-weights document AI and OCR pipelines
  • On-premise vision-language assistants
  • Chart and diagram understanding for analytics
  • Compliance and regulated-industry multimodal apps
  • Research baselines for vision-language models
08

Întrebări frecvente

Ce este Llama 3.2 90B Vision (multimodal)?

Llama 3.2 90B Vision (multimodal) este un model de Meta din categoria Multimodal. Este listat pe Railwail, dar nu poate fi rulat în acest moment.

Cât costă Llama 3.2 90B Vision (multimodal) pe Railwail?

Llama 3.2 90B Vision (multimodal) nu poate fi rulat pe Railwail în acest moment, deci nu există preț curent. Alternativele disponibile cu prețuri sunt listate mai jos pe această pagină.

Care este fereastra de context a Llama 3.2 90B Vision (multimodal)?

Fereastra de context a Llama 3.2 90B Vision (multimodal) conține 131.072 token-uri. Un răspuns poate fi lung de până la 8.192 token-uri.

Cât de rapid este Llama 3.2 90B Vision (multimodal)?

Nu sunt suficiente rulări măsurate ale Llama 3.2 90B Vision (multimodal) pe Railwail încă pentru a indica un timp de rulare. Depinde de intrare, de setări și de sarcina la furnizor.

Este Llama 3.2 90B Vision (multimodal) mai bun decât BLIP?

Depinde de sarcină. Llama 3.2 90B Vision (multimodal) (Meta) și BLIP (Salesforce) sunt ambele modele din categoria Multimodal. Pagina de comparație arată prețurile și specificațiile lor una lângă alta.

Compară Llama 3.2 90B Vision (multimodal) și BLIP

Poate Llama 3.2 90B Vision (multimodal) procesa imagini?

Da. Llama 3.2 90B Vision (multimodal) acceptă imagini ca intrare, pe lângă text.

Pot folosi Llama 3.2 90B Vision (multimodal) chiar acum?

Indisponibil în prezent: acest model a fost dezactivat. Pagina rămâne online; alternativele disponibile din aceeași categorie sunt listate mai jos.

Toate modelele printr-o singură API

O cheie API pentru fiecare model pe Railwail. Utilizarea se percepe din credite prepay, 1 credit = 0,01 USD.