Llama 3.2 90B Vision (multimodal)

MultimodalNon disponible
par MetaID du modèle: llama-3-2-90b-vision-mm

Meta's flagship vision-language model. 90B parameters, image understanding + chat, strong VQA performance.

Statut
Non disponible
Contexte
131 072 tokens
Max. sortie
8192 tokens
Entrée → Sortie
Texte + Image → Texte
Développeur
Meta
Mis à jour
25 juin 2026

Llama 3.2 90B Vision (multimodal) n'est actuellement pas disponible

Actuellement indisponible : ce modèle a été désactivé.

Vous pouvez toujours consulter les détails sur cette page. Choisissez l'une des alternatives disponibles ci-dessous pour exécuter immédiatement un modèle comparable.

Voir les alternatives
01

Modèles comparables

Tous dans cette catégorie
  • BLIPSalesforce

    Salesforce BLIP. Vision-language model for image captioning and visual question answering. Given an image it writes a short natural-language caption, or answers a question about the image when one is supplied. A widely used baseline for automatic captioning.

    ≈ 0,00030 $US/exécution

  • Anthropic's April 2026 flagship. 87.6% on SWE-bench Verified, 3x higher image resolution, output self-verification, vision + reasoning.

    6,00 $US/1M entrée

  • Anthropic's balanced mid-tier model from February 2026. Best price/performance for production workloads: 5x cheaper than Opus, near-flagship quality.

    3,60 $US/1M entrée

02

Playground

Essayer Llama 3.2 90B Vision (multimodal)

Chat

Actuellement indisponible

Actuellement indisponible : ce modèle a été désactivé.

Le terrain de jeu est désactivé. Vous pouvez trouver des modèles comparables dans la même catégorie : Parcourir les alternatives

Essayer Llama 3.2 90B Vision (multimodal)

Envoyez un message. La réponse arrive complète une fois que le modèle a terminé (sans streaming).

Prompt système
Longueur max. de la réponse (tokens)

Cette exécution

Pas de prix – actuellement indisponible.

Nouveau par ici ?

5 crédits gratuits (0,05 $US) à l'inscription avec Google

Utilisable 24 heures après l'inscription, jusqu'à 5 exécutions par jour et au maximum 2 crédits par exécution. Les autres méthodes de connexion commencent sans crédits.

03

À propos de Llama 3.2 90B Vision (multimodal)

RésuméAu 25 juin 2026

Llama 3.2 90B Vision (multimodal) est un modèle de Meta dans la catégorie Multimodal. Llama 3.2 90B Vision (multimodal) n'est actuellement pas disponible sur Railwail. La fenêtre de contexte contient 131 072 tokens, et une réponse peut faire jusqu'à 8192 tokens.

Arrière-plan

À propos de Meta AI (FAIR)

Fondée 2013 · Menlo Park, California, USA

Meta AI is the research arm of Meta Platforms, established in 2013 as Facebook AI Research (FAIR) by Yann LeCun. FAIR has open-sourced many foundational models including PyTorch, RoBERTa, DETR, SAM and the LLaMA family. LLaMA 1 was released in February 2023, LLaMA 2 in July 2023, LLaMA 3 in April 2024 and LLaMA 3.1 (405B) in July 2024. Llama 3.2 launched in September 2024 at Meta Connect, introducing the first multimodal models in the LLaMA family (vision-enabled 11B and 90B) together with tiny on-device text-only siblings (1B, 3B). All Llama 3.2 vision weights are released under the Llama 3 Community Licence and are widely used by enterprise customers via Meta's partner ecosystem (Hugging Face, AWS Bedrock, Azure AI Studio, Google Vertex, Together AI, Groq, Fireworks).

Visiter Meta AI (FAIR)

Architecture

Decoder-only Transformer with cross-attended vision encoder

Llama 3.2 90B Vision combines the 70B-parameter Llama 3.1 text backbone (extended to 90B with vision components) and a Vision Transformer image encoder integrated via cross-attention adapter layers, similar in spirit to Flamingo but reusing the LLaMA architecture. The vision tower processes each image to a sequence of visual tokens which are injected into specific cross-attention layers of the LLM decoder while the original text-only weights remain frozen during the multimodal training stage, preserving text-only performance. Pretraining used 6B image-text pairs followed by multi-stage supervised fine-tuning and Direct Preference Optimisation (DPO) on a curated set of image instructions, math and chart data. The model supports a 128K context window and accepts up to 1120x1120 image inputs natively (with tiling for larger images). It does not support video or audio. Llama 3.2 90B Vision is released under the Llama 3 Community Licence (free for commercial use under 700M MAU).

Paramètres
90B
Contexte
128 000 tokens

Capacités

  • Open-weights 90B vision-language model under Llama 3 Community Licence
  • 128K token context window
  • Image input up to 1120x1120 with tiling for larger images
  • Chart, diagram, OCR and document understanding
  • Strong on MMMU, MathVista, ChartQA and DocVQA among open-weights models
  • Multilingual: English, German, French, Italian, Portuguese, Spanish, Hindi, Thai
  • Tool use and JSON output via Llama 3.1 alignment recipe
  • Best for: open-weights multimodal apps, on-premise document AI, indie research

Entraînement et licence

Pretrained on 6B image-text pairs from public web and licensed sources; supervised fine-tuning and DPO on curated multimodal instruction data. Text knowledge inherited from Llama 3.1 (15T tokens).

Licence: Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.

Tests de sécurité: Meta publishes a comprehensive model card with red-team findings on CBRN, child-safety and hate-speech vectors, plus Llama Guard 3 and Prompt Guard 2 companion models for production safety.

Limitations connues

  • No video or audio input
  • Latency and cost dominated by 90B params; requires multi-GPU serving
  • Licence restricts the largest hyperscaler use cases
  • Vision quality below GPT-4o and Claude 3.5 Sonnet on hardest charts
  • English-centric in vision domain
04

Tarification

Actuellement indisponible : ce modèle a été désactivé. Il n'y a pas de prix pour ce modèle pour le moment, il ne peut donc pas être exécuté.

05

API

Appelez Llama 3.2 90B Vision (multimodal) avec votre clé API Railwail. Utilisez cet ID de modèle dans la requête :

Actuellement indisponible

Le modèle n'a pas de prix vérifié ou est désactivé ; les appels API sont refusés.

06

Spécifications

ID du modèle
llama-3-2-90b-vision-mm
Développeur
Meta
Catégorie
Multimodal
Entrée
Texte, Image
Sortie
Texte
Fenêtre de contexte
131 072 tokens
Sortie max.
8192 tokens
Taille du modèle
90B
Licence
Llama 3 Community Licence: free for commercial use up to 700M MAU; redistribution must include the licence and acceptable use policy.
Entrée du catalogue mise à jour
25 juin 2026

Étiquettes

  • meta
  • llama
  • multimodal
  • vision
  • open-weights
07

Cas d'usage

À quoi ça sert

  • Open-weights document AI and OCR pipelines
  • On-premise vision-language assistants
  • Chart and diagram understanding for analytics
  • Compliance and regulated-industry multimodal apps
  • Research baselines for vision-language models
08

Questions fréquemment posées

Qu'est-ce que Llama 3.2 90B Vision (multimodal) ?

Llama 3.2 90B Vision (multimodal) est un modèle de Meta dans la catégorie Multimodal. Il est listé sur Railwail mais ne peut pas être exécuté pour le moment.

Combien coûte Llama 3.2 90B Vision (multimodal) sur Railwail ?

Llama 3.2 90B Vision (multimodal) ne peut pas être exécuté sur Railwail pour le moment, il n'y a donc pas de prix actuel. Les alternatives disponibles avec leurs prix sont listées plus bas sur cette page.

Quelle est la fenêtre de contexte de Llama 3.2 90B Vision (multimodal) ?

La fenêtre de contexte de Llama 3.2 90B Vision (multimodal) contient 131 072 tokens. Une réponse peut faire jusqu'à 8192 tokens.

Quelle est la vitesse de Llama 3.2 90B Vision (multimodal) ?

Il n'y a pas encore assez d'exécutions mesurées de Llama 3.2 90B Vision (multimodal) sur Railwail pour indiquer un temps d'exécution. Cela dépend de l'entrée, des paramètres et de la charge chez le fournisseur.

Llama 3.2 90B Vision (multimodal) est-il meilleur que BLIP ?

Cela dépend de la tâche. Llama 3.2 90B Vision (multimodal) (Meta) et BLIP (Salesforce) sont tous deux des modèles de la catégorie Multimodal. La page de comparaison affiche leurs prix et spécifications côte à côte.

Comparer Llama 3.2 90B Vision (multimodal) et BLIP

Llama 3.2 90B Vision (multimodal) peut-il traiter des images ?

Oui. Llama 3.2 90B Vision (multimodal) accepte les images en entrée en plus du texte.

Puis-je utiliser Llama 3.2 90B Vision (multimodal) maintenant ?

Actuellement indisponible : ce modèle a été désactivé. La page reste en ligne ; les alternatives disponibles de la même catégorie sont listées plus bas.

Tous les modèles via une API

Une clé API pour tous les modèles sur Railwail. L'utilisation est facturée à partir de crédits prépayés, 1 crédit = 0,01 $US.